注册并分享邀请链接,可获得视频播放与邀请奖励。

与「ReinforcementLearning」相关的搜索结果

ReinforcementLearning 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 ReinforcementLearning 的内容
The future of humanoids isn't just mechanical—it's neural. A humanoid robot is an embodied AI system. Motors, sensors, and actuators provide the body, but neural intelligence provides the mind. Neural models power vision, speech, language, navigation, manipulation, planning, memory, decision-making, and continuous learning. As robotics evolves, nearly every cognitive subsystem is becoming neural-first. That's why represents more than a niche—it's a foundational concept for the next generation of intelligent robotics. #HumanoidRobots# #EmbodiedAI# #PhysicalAI# #NeuralNetworks# #ArtificialIntelligence# #Robotics# #MachineLearning# #DeepLearning# #RobotLearning# #ReinforcementLearning# #ComputerVision# #GenerativeAI# #AIResearch# #Automation# #FutureOfAI#
显示更多
🧵 Deli AutoResearch SKILL is now officially open source! 🎉 Alongside it, we’re dropping our 4th survey paper — this time on Self-play. Inspired by AlphaZero, we got a powerful insight: prior knowledge doesn’t always lift the ceiling. Models can discover more globally optimal solutions just by playing against themselves. The biggest change in this paper? For the first time, the AutoResearch Agent autonomously planned GPU experiments — and submitted actual RL runs on the DeepSeek 285B model. The entire RL pipeline — experiment design, code writing, running, debugging, and conclusion summarization — was 100% automated, with zero human intervention from me. This was incredibly difficult, but an incredibly important step. GRPO is the tool being called by the AutoResearch Agent here. We see this as the beginning of our Continual Learning research journey. 🚀 As always, this is my personal research project, unaffiliated with any organization. All views are my own. #AI# #ReinforcementLearning# #SelfPlay# #OpenSource# #AutoML# #ContinualLearning# #DeepSeek#
显示更多
0
15
1.1K
168
转发到社区
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:
显示更多
0
352
8.7K
857
转发到社区
BIG ANNOUNCEMENT FROM HUGGING FACE TODAY: We're unveiling Microduck 🐥🤖 It's a tiny $399 open-source robot you can teach new tricks with reinforcement learning. It can walk, pick things up, get back up when it falls, and even roller-skate. Welcome to the era of open-source affordable robots to democratize physical AI and world models! 🤗🤗🤗
显示更多
0
546
11K
832
转发到社区
We have a huge news to share today! Today we are unveiling the first truly accessible RL robot - welcome Microduck A 25 cm tiny open-source biped with 15 actuators and packed with sensors (camera, speaker, LiDAR, NFC, bluetooth, wifi, etc) that you train yourself with reinforcement learning. It's also playable out of the box with more than half a dozen fun and playful pre-trained policies to have it walk, sit, crouch, roller-skate, pick up objects with its articulated beak, and recover on its own. And all for less than $400. See all the details, play with the simulator and order it at: (video with sound on 🔊)
显示更多
0
499
6.8K
739
转发到社区
Instead of generating a finished image, this AI paints by writing editable JavaScript. The project uses reinforcement learning and hand-rated examples to teach Qwen 3.5 how to make a good watercolor.
显示更多
0
40
1.8K
86
转发到社区
Hi! Recapping some changes we have rolled out over the last couple of weeks that have further reduced the risk associated to potentially destructive actions being performed by Codex during its work. A few weeks ago, we started investigating a small number of reports where GPT-5.6 in Codex took destructive actions outside what the user asked for. The most serious pattern we found was a command meant to clean up temporary work that could instead delete the user files. This should obviously not happen. Here’s what we found: - Codex sometimes creates temporary folders while working and cleans them up afterward. In rare cases, GPT-5.6 got that cleanup wrong. One pattern involved reusing a system environment variable like $HOME for temporary work. A malformed cleanup command could then point at the actual home directory instead of the temporary folder. - There were cases where the model tried to delete or overwrite a temporary path without checking what was already there. We’ve added protections at several layers: - Codex is now explicitly instructed to check deletion targets before acting, create fresh temporary directories, avoid repurposing system environment variables, prefer recoverable actions, and stop when the scope is unclear. - We strengthened the execution checks that identify high-risk deletion commands and escalate them for review. If a command is rejected, the model is directed to take a safer approach. - We made Full access harder to enable accidentally, added clearer warnings, and further restricted especially risky permission combinations. - We updated Auto-review to better identify destructive actions. - We built targeted evaluations that replay the failures we observed. We’re also adding reinforcement-learning tasks and graders focused on these risks, and filtering destructive actions from training data. In those replay evaluations, the changes substantially reduced the behavior while preserving Codex’s ability to complete normal coding work. Two things to do on your end: - Keep the Codex app up to date. We are always improving safety, performance and many other things. - Use one of the sandbox modes: "Ask for approval" or "Approve for me". Only use Full access for environments you trust and can recover. Thanks and happy Codexing out there!
显示更多
0
361
1.2K
74
转发到社区
A humanoid robot is fundamentally an embodied AI system. While motors, actuators, batteries, and sensors provide the body, neural intelligence provides the mind. Without neural computation, a humanoid is simply an advanced machine. With it, the robot becomes capable of adapting to unpredictable environments, understanding language, recognizing objects, planning tasks, and continuously improving through experience. This shift positions "Humanoid Neural" as a foundational concept within the robotics ecosystem. Neural Intelligence is the Core of Future Humanoids. Modern humanoids rely on neural-network-based systems to perform nearly every cognitive function. These include: Visual perception Speech recognition Language understanding Object identification Human pose estimation Motion planning Reinforcement learning Dexterous manipulation Long-term memory Decision making Navigation Emotional recognition Social interaction As robots become more capable, nearly every subsystem transitions from traditional programming toward learned neural models. This makes "neural" less of a niche AI term and more of an umbrella for robot intelligence. A Broad and Scalable Brand One of the strongest characteristics of is that it is not confined to a single product category. #Languageunderstanding# #neuralnetwork# #socialinteraction# #neuralntelligence# #smarthumanoids# #iq#
显示更多
Blockchain Sports Drift — one year later One year ago, an AI drift car was just an idea. Today, it's real — and it drives itself. Here's what we've built so far. The car drives itself Our prototype drives completely on its own. No driver. No remote control. The AI reads the car's position, speed, and angle 30 times every second — and acts on it. It doesn't just react. It acts. The AI keeps learning Instead of following the same instructions every time, the system learns through reinforcement learning. It practices thousands of times, learns from every run, and keeps getting better. Always knows where it is The car combines information from several sensors into one accurate picture of where it is on the track. Even while sliding sideways, it always knows its position. Safety comes first The car has six independent hardware safety systems. They don't rely on software or a radio connection. If something goes wrong, the car stops automatically. We've already demonstrated this during live events with no safety driver inside. Building a unique dataset We're collecting biometric data from professional drivers using our own specialized racing suit. Combined with everything the car records while driving, we're building a unique dataset for autonomous motorsport. Visitors from around the world Over the past year, dozens of guests from the United States, Japan, and across Europe visited our facility. They saw the technology, tested the car, and experienced our professional drift cars firsthand. The reaction was the same every time — disbelief, then excitement. What’s next We’re building the foundation for a new kind of motorsport. First, we’ll prove what our own AI can do. Then we’ll open our core technology to the world’s leading tech companies — so they can build their own AI-powered drift cars and compete. The destination? A world championship in the United States — where AI cars backed by the world’s biggest tech companies compete against the best human drift drivers. All powered by our core technology. This is only year one!
显示更多
After Xiao Wu, meet Xiao Liu. 🤖 Tencent Robotics X’s latest robot isn’t just learning massage moves. It’s learning the touch behind them — where to press and exactly how hard. A custom data system captures vision, touch, and force from human demonstrations. Using reinforcement learning, Xiao Liu learns to reproduce both the trajectory and the pressure. Getting to the right spot is one thing. Getting the pressure right is another. My shoulders volunteer as tribute.
显示更多