
SO-100 අතිරේකයේ එක් human-gated DAgger වටයක පියවරෙන් පියවර විස්තරය: policy ධාවනය කරන්න සහ එය වාර්තා කරන්න, එය වැරදි වූ විට පිළිගන්න, ධාවනය නිරෝධනයක් ලෙස ගොනුගත කරන්න, මිශ්ර දත්තකට සෑදින්න, සහ checkpoint එකක් සිට පුහුණුවීම දිගටම කරන්න. නායක බාහුවක් නොමැති පිරිසට සඳහා keyboard සහ slider takeover පතුරුම්ය ඇතුළත් කරයි, සහ වටයක් අර්ථවත් නොවන කරන හතර නිරුපිතක් ඔබ දනිය යුතුය.
ඔබගේ policy ධාවනය වේ. එය ඝනකයට ඉක්මවා යි, gripper සෙන්ටිමීටරයක් ඉතා කලින් එය වසා ගනී, සහ එය එතනින් ඉහළට යි. කිසිවක් දෝෂයට පත් නොවන, සහ පුහුණුවීම හා දෑත් නිරීක්ෂණයෙන් එය පැහැදිලි නොවේ. නිවැරඳුම ඉතිරි දෙයක් හෝ 20 000 gradient පියවර හෙට දිගටම අධ්යයන ප්රදර්ශන වලින් පැමිණවීම නොවේ. එය ඔබගේ අත සිටින ස්ථානයට නැවතත් ඇතිරේකයට ස්ථාපනය කරන්න එහි වැරදි ස්ථානයේ, එයින් ඔබ කළ දෙයද්ර වාර්තා කරන්න, සහ ඉතිරි දත්ත සහ එම නිවැරඳුම මත ඉදිරි checkpoint පුහුණුවීම කරන්න. එය DAgger වටයක් වන අතර මෙය එකක් SO-100 සහ vision-language-action policy සමඟ ධාවනය කරන ආකාරයයි.
සිද්ධාන්තය වෙනත් තැනක ඇත: dataset aggregation ක්රියා කරන්නේ මන්ද සහ human gating එය සම්බන්ධයෙන් සිට වෙනස් කරන දෙය. මෙය ක්රියාකාරි මෙහෙයුම් නිෂ්පාදනාකාරය, සහ එය පුහුණුවිත checkpoint එකක්, ක්රියාකාරි කැමරා කට එකක්, සහ චලනය වන අතිරේකයක් උපකල්පනය කරයි. පහත සඳහන් පියවර හයක් වටය ලෙස මෙම පැතිවල DAgger පිටුව මත ක්රියාත්මක වේ, නමුත් අනුක්රමය ඔබගේම ස්ක්රිප්ට් එකිනි.
වටයක් කෙටියෙන්
- •පුහුණුවිත policy ධාවනය කරන්න සහ එය වාර්තා කරන්න, ධාවනයේ කාර්ය පෙළ සහිතව සාමාන්ය දුරස්ථ ලේබල් වෙනුවට.
- •අතිරේකයේ තත්පරය තුරා පිළිගන්න: නායක අතිරේකයක් සහිත mirror handover එක, keyboard හෝ sliders නැතිවම ක්ෂණිකව සහ දෙමුහුන්.
- •සෑම ධාවනයම කඩුටු කරන්න: එය නිවැරඳුම ලෙස ගොනුගත කරන්න, එය evaluation episode එකක් ලෙස තබා ගන්න, හෝ එය ප්රතික්ෂේප කරන්න.
- •සකස්කිරීම අතින් මිශ්ර කරන්න - මුල් පෙන්වීම් සහ නිවැරඳුම්, එපිසෝඩ් ප්රභවය අනුව තෝරා ගන්න. කිසිවිට නිවැරඳුම් පමණක් පුහුණුවීම නොකරන්න.
- •අවසාන checkpoint එකෙන් පුහුණුවීම දිගටම කරන්න, සහ checkpoint එක කුමන මිශ්ර මිශ්ර එක නිෂ්පාදනය කළ බව සටහන් කරන්න.
- •වටයට කිසිවක් දිරිමත් කිරීම ඉතිරි අධ්යයන දෑත් නිරීක්ෂණයට වඩා intervention rate ගිණුම යි.
දෙවන වටය ඉතිරිව තරමක් දත්තක් පමණක් නොවේ
Behaviour cloning එක පුහුණුවා ගරු කරන්න ඉතිරි ලෙස දිස්තු වූ ප්රකාශ නොවේ. පරීක්ෂණ අවස්ථාවෙහි policy නිෂ්පාදනය කරන ඉතිරි ලෙස දිස්තු වූ ප්රකාශ දර්ශනයට ගිය විට, සහ කුඩා ක්රිය දෝෂ ඉතිරි ලෙස දිස්තු නොවූ ප්රකාශ ස්ථාන ඇතුළත් කරයි. Ross, Gordon සහ Bagnell එම ඉතිරි දෑත් AISTATS 2011 සඳහා නිර්වචනය කර, එය ස්ථාවර තීරණාත්මක policy එකක් සහ ඔවුන්ගේ අඩු කිරීම, තෝරාගෙන ඔබ පැමිණවීම සිට ඉතිරි ස්ථාන සඳහා පුහුණුවීම ඇත විසඳුවා ඇත: වත්තමාන policy ධාවනය කරන්න, විශේෂඥ තෝරාගෙන ඔබ පැමිණවීම ස්ථාන සඳහා ඉතිරි ලෙස දිස්තු වූ, එය ඉතිරි ලෙස දිස්තු කරන්න, පුනරාවර්තනය කරන්න. Kelly et al. HG-DAgger සහිතව විමසීම ප්රයෝගිකව සිදු කර, එහිදී මිනිසු සම්පුර්ණ පිහිටුවීම සහ හෙවත් තිරැවණ ලබා ගැනීමට, කිසිවක් පිහිටුවීම ගුණ දිගටම කරන්න, සහ simulated සහ සැබෑ autonomous driving කාර්ය මත behaviour cloning සහ DAgger දෙකම සඳහා වඩා හොඳ ක්රියාකාරිතාවය වාර්තා කරයි. Human gating එක වටයක් desk arm එක මත tolerant කරන දෙයක් ලෙස සිටුවීම ලබා දෙයි - ඔබ ඔබගේ අත චලනය කරනු ඇත කිසිවක් වැරදි වූ විටම පමණක්.
දෙස්තරයක් සිතුම් සඳහා වැදගත් වේ: නිවැරඳුම් සාමාන්ය පෙන්වීම් නොවේ: ඔවුන් Mandlekar et al. විස්තර කරන බොතල්නෙක් ස්ථාන මත හේතු කරයි, එහිදී කුඩා අපගම්යතා policy එකට ස්ථාන පුරුදු නොවූ තුළට හරිහැටි ගිරිසර කරයි. සහ dataset එක නිරුවිතින්න ඉතිරි දෛඩ උපකිරීම වලින් බැසට එකක් එක්ස්ට බදින, නමුත් යුතු ලෙස සිතුම් ප්රකාශ රටාවෙනි - Belkhale, Cui සහ Sadigh තර්ක කරයි දත්ත දෙස සිට තර්ක ස්ථාන විවිධතා සෑම විටම ප්රයෝජනවත් නොවන, සහ ක්රිය අපගම්යතා සහ සංක්රමණ විවිධතා එක්ගෙන dataset ගුණ තීරණය කරයි. Mixed dataset එක සම්මතිකිරීම එකක්, එය වලතුලින්.
වටයක් පටන්ගැනීම සිට පෙර මෙයි අසරණ කරන්න
DAgger වටයක් policy එකක් තුලනය කරයි ඉතිරි තුලනයට පුරා. කිසිවක් ඔබ වටයතුලින් සිට වෙනස් කරන dataset එක නොවේ එම සම්මතිකිරීම තුලතින්.
- කැමරා පිහිටුම් සහ ගුණුවුම්, wrist camera එක ඇතුළත්. ඔබ ප්රතිසන්ධිකිරීම එකක් ලොස් කරන්න සහ ඔබ නිරීක්ෂණ බෙදාගෙන ගිස්ගිස්කිරීම නිසා, policy එක නොවේ.
- Exposure සහ white balance, ඔබගේ capture stack එක ඔබට ඒවා pin කිරීමට ඉඩ දුන්නේ නම්. Auto-exposure drifting වටයතුලින් domain shift එකක් අගෝපනීයතා දිගටම කිරීම ලෙස සිටුවීම ලබා දෙයි.
- Arm calibration සහ servo zero පිහිටුම්. ඔබ recalibrate කිරීමට එක්ස්ට තිබිය යුතුයි, එය පෙර සෙමින්තත් වාර්තා කර dataset එකක් තෙහි ලෙස ගණනය කරන්න.
- කාර්ය පෙළ. සෑම VLA තුලතින් එය උපකර්තනය කරයි; එය සිට mid-loop එක ගෙයින්දිස්කිරීම එක කර්තව්ය ලෙස සිටුවීම ලබා දෙයි.
- Lighting, table surface, object set. නව object එක නව පරීක්ෂණයක් ලෙස සිටුවීම ලබා දෙයි, තිඩි වටයක් නොවේ.
- වාර්තා කිරීම frame rate එක. Intervention rates තුලනය කිරීම දෙස sampling rasters සිට දෑත් ගිණුම එකතු කිරීම ගිණුම සිට ගිණුම ලබා දෙයි.
Hsu et al. අතිරේකයේ සිත-centric දෑත් සිට තුලනය කර සම්මතිකිරීම තුලතින්-දසෂ්ටි දෑත් සිට ඉතිරි ඉතිරි විස්තර සඳහා සඳහා විකලිතිරම නිරීක්ෂණ බෙදාගෙන ගිස්ගිස්කිරීම අවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුවිසිඳුවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුමිතිරි ගිණුම සිට අවිසිඳුවිසිඳුවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුවිසිඳුවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුමිතිරි, නිසා දිස්තු අඩු ගිණුම සිටුම ඇතිකිරීම. Five-joint arm එක මත, gripper timing එක සාමාන්යයෙන් ඔබගේ නිවැරඳුම් නිසා සිටුවීම ලබා දෙයි සිතුම් නිසා, සහ gripper timing එක wrist දෑත් සිට වටි.
වටයක් තුලතින්
- 1ඉන්දජනය ධාවනය කරන්න සහ එය වාර්තා කරන්න
පූර්වලිපි checkpoint එක විරුද්ධ ධාවනය සිට පටන්ගැනීම, ඉන්දජනය නිරීක්ෂණ root තුළ වාර්තා කිරීම සිට සිට පටන්ගැනීම. එම ඉතිරි එක ධාවනයේ නිරැඔබ පෙළ උත්තරාධිකාර සිට එක, නිසා සිටුවීම ලබා දෙයි දුරස්ථ ලේබල් එක පෙර default සිට. වාර්තා කිරීම නැතිවම ඔබ මුහුණු අසफලතා දෙස පිටුවට පරීක්ෂණ කරණු ඇත නිසා එය පුහුණුවීම ගිණුම නිසා.
bash# two calls, not one: the run, then its recording POST /inference/start # model_id, and hf_repo_id = the checkpoint to drive POST /recording/start # root=inference # root=inference also makes the recording inherit the run's task text - 2පිළිගන්න එය වැරදි වූ විටම
Take over හෝ ඉතිරි ඉතිරි input mode එක තෝරා ගන්න: නායක අතිරේකයක්, keyboard හෝ sliders. Runner එක නතර කරනු ලැබේ, ඔබ නිවැරඳුම්, ඔබ ඉතිරි ලෙස සිටුවීම ලබා දෙයි. Frames වාර්තා කිරීම ඔබ ධාවනය කිරීම එ้ස්ස්ව සිතුම්ව้ Frames වාර්තා කිරීම ඔබ ධාවනය කිරීම එ้ස්ස්ව සිතුම්වසිටුවීම ලබා දෙයි.
bashPOST /inference/takeover/start # input = leader | keyboard | sliders POST /inference/takeover/nudge # keyboard, relative delta per call POST /inference/takeover/set # sliders, absolute target POST /inference/takeover/stop # back to the policy - 3Episode එක තුලතින් කඩුටු කරන්න
තීරණය ඉතිරි episode එකක් සිට: නිවැරඳුම් ලෙස ගොනුගත කරන්න, evaluation ලෙස තබා ගන්න, හෝ ප්රතික්ෂේප කරන්න. ධාවනය policy එක පෙර ඉතිරි උපකාරය නැතිවම ඉතිරි කිරීම evaluation දත්ත.
- 4නිවැරඳුම් dataset එක සුසම කරන්න
Corrections එක සংගතුයි local dataset එකක් ඉතිරි policy සිට සහ cloud storage එක තුළ ස්වයංක්රිය sunc එකෙන්. කිසිවක් මිශ්ර සිටුවීම ලබා දෙයි ඔබ එහි තබා ගන්නේ නම්.
- 5Mixed dataset එක සෑදින්න
Combine එක original dataset එක නිවැරඳුම් dataset එක සඳහා, episode තෝරා ගන්න explicitly ඉතිරි ප්රභවය සිට. ඉතිරි ඉතිරි එක ordinary
bashPOST /training/datasets/compose sources = [ original_dataset, korrekturen_<policy> ] episodes = explicit selection per source - 6LeRobot dataset
එක එකතු කිරීම ඉතිරි පුහුණුවීම පිළිබඳ.
bash# field on the training job base_checkpoint = s3://ay-robots/checkpoints/<run>/<checkpoint> # the platform passes it to the training pod as BASE_CKPT_S3
Checkpoint එකෙන් පුහුණුවීම දිගටම කරන්න
Mized එක පුහුණුවා ගරු කරන්න පෙර checkpoint සිට නිසා base model එක සිට. සටහන් checkpoint සහ mized එක; එම ගණුවුන් නැතිවම වටයක් reproducible නොවේ.
පියවර 2 තුලතින්: පිළිගැනීම ක්රම දෙසනායක අතිරේකයක්එ
leader-follower
mode තුල පිළිගැනීම handover එක දෙස අතිරේකයක් නිසා පිහිටුම ඉතිරි එකක් සිට. Take over හෝ ඉතිරි ඉතිරි runner එක නතර කරනු ලැබේ සහ driver කර නායක දිස්තු රැරැ follower දිස්තු pose එක, සිට කිසිවක් jump සිට torque ඉතිරි ගිණුම. එම පිහිටුම alignment drive එක timeout සිට, ඔබ align නායක අතින් සහ release පිහිටුවීම ලබා දෙයි පෙර දෙස පරාසය සිට පහ අංශක එකතු. ඉතිරි ඔබ teleoperate සාමාන්යයෙන් සහ ක්රිය column එක වාර්තා ඔබ සිතුම්.
Honest සිතුම් දෑත් පතුරුම්ය: alignment drive එක සහ torque handover එක සිටුවීම ලබා දෙයි least tested පිහිටුවීම වටයක් real hardware එක. පරීක්ෂණ handover එක slow, සිතුවීම pose එක පෙර ඔබ rely සිට වටයක් ඔබ care සිට. Nayak අතිරේකයක් produces smoothest නිවැරඳුම් තුලතින් ක්රම තුලතින්, සහ තිබිය හැකි වඩාත් එකතු කිරීම mechanics සිට.
| නායක අතිරේකයක් නැතිවම: keyboard සහ sliders | බහුතරයක් පිරිසු කියවීම සිට නිසා තිබිය එක අතිරේකයක්. එය sufficient. තෝරා ගන්න keyboard හෝ slider input එක පිළිගැනීම ඉතිරි පිහිටුවීම ලබා දෙයි, සහ පිළිගැනීම සිට ක්ෂණිකව සහ දෙමුහුන් - තිබිය කිසිවක් දෙවන අතිරේකයක් සිට align, සිට තිබිය alignment පියවර. Follower එක තබා ගන්න පිහිටුම සිට සහ බලා ගන්න input සිට. | Input mode එක | Arm එක චලනය කරන ඉතිරි |
|---|---|---|---|
| Per-call limit enforced එක server එකෙන් | Locked ඉතිරි | Nayak අතිරේකයක් | Mirror එක driver follower එක nayak එකෙන් joint angles එකතු |
| කිසිවක් nudge හෝ set calls දෙස mode එක; mirror එක writes follower goals ක්ෂණිකව | Never locked, සහ default පිහිටුවීම ලබා දෙයි input mode එක ඉතිරි එකතු නම් - නිසා එය අවශ්ය දෙවන අතිරේකයක්; නිසා নායක id එක පිළිගැනීම සිට refused | Keyboard එක | Relative nudge එක per key press එක, sent සිට takeover nudge endpoint එක |
| Hard clamp එක 2 degrees එක per joint එක, 4 degrees එක gripper සිට | Rejected එක 409 එකතු පිහිටුවීම ලබා දෙයි takeover එක started නැතිවම nayak mode එක | Sliders එක | Absolute target pose එක, sent සිට takeover set endpoint එක |
At most 6 degrees එක travel එක toward target එක per call එක; interface එක keeps sending about ten times එක තේ
Q / A joint 1 R / F joint 4
W / S joint 2 T / G joint 5
E / D joint 3 Z / X gripper
Training guide දැක්කෙ ordered පියවර එ GR00T fine-tuning ධාවනය එ SO-100 dataset එක
Training පිහිටුවීම ලබා දෙයි එ සිට same guided sequence එ පටල ධාවනය එක; checkpoint පිහිටුවීම ලබා දෙයි field එක differs එක.
Early එ්ස්ලැබ්ඩු. Correction එක starts පසු gripper එක closed එ nothing එ teaches recovery එ්ස්ලැබ්ඩු failure එ policy එක should තිබිය නම්ට entered, සහ recovery දත්ත එე worth අඩු තරම් avoidance දත්ත එක. Interrupt එ්ස්ලැබ්ඩු පටල moment එ ඔබ confident එ trajectory එ wrong එක, correct එ්ස්ලැබ්ඩු difficult පිහිටුවීම ලබා දෙයි, hand එ්ස්ලැබ්ඩු back එ්ස්ලැබ්ඩු soon එ state එ නිසා handled පෙර policy එක. ThriftyDAgger එ automated එ තීරණය එ gating එ interventions එ novelty එ estimated එ risk එ under එ fixed එ human එ budget එ, නිසා එ single එ arm එ සිට එ human එ already එ watching එ, human එ gate එ cheaper එ සහ better එ calibrated එ තරම් එ anything එ ඔබ tune එ.
- හසිට අතින්
- Doable එ්ස්ලැබ්ඩු open-source එ stack එ සහ කිහිපය එ scripts එක. එ costs එ නිසා bookkeeping එක, සහ bookkeeping එ นිසා DAgger එ rounds එ die එක.
- Write එ frames එ ඔබගේ පෙනුම inference එ script එ කිරීම නිසා LeRobot එ dataset එක, සිට task එe string එ policy එ trained එ.
- Pause එ policy එ loop එ, switch එ command එ source එк, සහ flag එ තරම් frame එ ඔබ drive එ නිසා intervention එk. Nැතිවම flag එk, corrections එ නිසා ordinary එ demonstrations එk.
- Decide එ deliberately එ එ happens එ තුලින් transition එe frames එ පිටුවට policy එe releasing එe control එe සහ ඔබගේ පටල input එk.
Point එe fine-tuning එe entry එe point එe අතින් පෙර checkpoint එk, සහ check එe එ log එ එ loaded එe එ weights එk.
Over එe platform එkසිට සිට එ buttons එk. එ automated එ නිසා එ easy එ නිසා wrong එe අතින්: එe per-frame එe intervention එe flag එk, එe split එe පිටුවට corrections එk සහ evaluations එk, සහ record එe නිසා checkpoint එk produced එe එ mix එk. Nැතිවම enters එ composed එe dataset එe එ ඔබ did එe select එk.එ does නැතිවම decide එe ඔබ සිට. එ run එe counts එe නිසා correction එk, එe episodes එe නිසා mix එk, සහ ඉතිරි intervention එk rate එk එ stopped එe falling එk remain එe judgement එe calls එk. එ fields එe documented එe under training, එe input එe modes එe under
teleoperation
.
- පියවර 3 තුලතින්: triage බලා ගන්න ගුණ
- පසු දිගටම ඔබ තිබිය එ recording එe සිට තරම් frames එe marked එe නිසා interventions එk. තුල destinations එe exist, සහ එ wrong එe එකතු quietly එe poisons එe පටල round එk.
- File නිසා correction එe ඉතිරි intervention එe تنیయ එ fix: එ policy එe තිබිය දිස්තු somewhere එe wrong එk සහ ඔබගේ input එe showed එe එ right එe තරම් එ state එe එ policy එe තිබිය produced එk.
Discard එe runs එe ruined එe අතින් එe unrelated එe - එ dropped එ camera එ frame එک, එ stalled එ servo එک, එ object එ ඔබ knocked එ over එک. Messy එe correction එe නිසා worse එe තරම් නැතිවම correction එk.
Freeze එe frames එk belong එk එ raw එe recording එk, නැතිවම එ training එ දත්ත
පිටුවට එ policy එ releasing එ control එ සහ ඔබගේ පටල input එ, එ arm එ holds එ still එ එ recorder එ keeps එ writing එ - එ run එ නිසා identical එ poses එ paired එ සිට slightly එ different එ images එk. Here එ handover එe frames එk stay එk එa raw එe recording එk සහ out එ නිසා correction එ dataset එk. If එ ඔබ build එ loop එe තිබිය, cut එe ඔවුන් deliberately: එe policy එe trained එe නිසා ඔවුන් learns එe නිසා pause එe නිසා ඔබ should එe act එk.පියවර 5 තුලතින්: composing එ mix එkComposition එk takes එk එ original එk dataset එk plus එa correction එk dataset එk සහ produces එk එ new එk, ordinary
LeRobot dataset

Dataset එa recording එa view එa showing එa episodes එa නිසා SO-100 එa dataset එa සිට camera එa streams එa
Correction එa episodes එa නිසා ordinary එa episodes එa සිට එ per-frame එe intervention එe flag එk, සිට ඔවුන් compose එk සිට එ original එk dataset එk nැතිවම conversion එk.පියවර 6 තුලතින්: එ continuing එe නිසා checkpoint එe really එe means එkTraining එe එ mix එe නිසා base එe model එe works එk නිසා throws එk away එk එ පෙර round එk සහ costs එk පූර්ණ ධාවනය. Continuing එe නිසා පෙර
එ faster එa සහ usually එa better එa. එ also තිබිය එa limited එa තරම් එe phrase එe suggests එk.
| Weight එe initialisation එe නිසා නැතිවම optimizer එe resume එe | පූර්ණ weights එe checkpoint එe contains එe එ parameters එe සහ nැතිවම else එk. Loading එe එ gives එe එ පටල ධාවනය එ better එa starting එa point එe තරම් base එe model එk, නිසා optimizer එk moments එk, learning-rate එe schedule එe position එe සහ දත්ත ए order එ සිට start එe නිසා zero එk. Expect එa එ loss එe spike එe එ බිම එe නිසා පටල ධාවනය, දෙස නැතිවම read එე නිසා failure එک, සහ දෙස නැතිවම call එa එ round එe نिसए resume එe. එ නිසා warm එe start එk. | Policy | Size | GPU tier | Inference per action step |
|---|---|---|---|---|---|
| Dataset format | Episodes before it is worth trying | GR00T N1.7 | about 3 B, roughly 40 M trained during fine-tuning | A100 80 GB or H100 80 GB | about 152 ms |
| LeRobot v2.0 or v2.1 | 50 | GR00T N1.5 | about 3 B | A100 80 GB or H100 80 GB | about 165 ms |
| LeRobot v2.0 or v2.1 | 50 | Pi0.5 | about 3 B on a PaliGemma backbone | A100 80 GB or H100 80 GB | about 485 ms |
| LeRobot v3.0 | 50 | SmolVLA | about 450 M | RTX 4090 or any 24 GB card | about 245 ms |
| LeRobot v3.0 | 30 | ACT | about 80 M, trained from scratch | RTX 4090 or any 24 GB card | about 20 ms |
LeRobot v3.0
50Latency එa compounds එa එa DAgger එa loop එa එ නිසා එ දෛඩුම්: එa roughly එa 485 එa ms එa per එa action එa step එa ඔබ take එa නිසා නිසා එ arm එa hesitated එa, නිසා එ wrong එa, සහ hesitation එe නිවැරඳුම් නිසා නැතිවම useful එ training එ දත්ත එk. If එ ඔබ repeating එa නිසා දත්ත එ එ නිසා chasing එa පූර්ණ success එ rate එe, repeat තුල එ fast එe model එk. Shukor et al. describe එe SmolVLA as එa designed එe នિසา train එe එa single එa GPU සහ deploy එa නිසා consumer එe GPUs සිට CPUs එe, සිට පුටුවතු inference එk stack එk එa decouples එk action එk prediction එk නිසා execution េnिसა allow එa higher එa control assertEquals - එe property එe නිසා keep එe එ takeover එe loop එk responsive ए.එ dataset එe formats එe නිසා නැතිවම interchangeable එ either എ. GR00T එე takes එe LeRobot එe v2.0 සිට v2.1 එe, සහ එ Isaac-GR00T repository එa describes එa එ input එa නිසා එ flavour එe නිසා LeRobot එe v2 format එe සිට එ added එ modality එe description එe file එe; එ newer එa trainers එe here එe expect එe v3.0 එe. එ mix එe composed එe එ wrong එe version එe fails එe එ load එe time එa එ producing තරම් bad එe policy එ - එ better එ failure එe mode එe, still එa එa wasted එe queue එe slot එe. එ
dataset එ documentation එe
lists එa එ format එa තर॑ trainer එe takes එk.
එ loop එe, සිට එ bookkeeping එe already එa done එaTakeover එa සිට නචක අතිරේකයක් සිට keyboard සිට sliders එe; per-frame إ intervention එe marking එe; filing එa runs එe නිසා corrections එa සිට evaluations එe; composing එe එ mixed එa dataset එe សිට එ explicit එe episode එe selection එa per එa source එe; සහ continuing え training එa నిสา checkpoint එe තුල base එa model එe. එ stays එa ඔබගේ තීරණය එe නිසා එ run එe counts എ nිසা correction එk, එ නිසා mix එe, සහ ඉතිරි intervention එk rate එa එ stopped එe falling එa.
See েnিසా එ DAgger එ loop එე wired
ඔබ wasting එ round එ ක්රම හතර
1. Training නිසා එ corrections එe alone
එ බොහොම පොදු failure එ සහ බොහොම ඉසිකිරීමේ shortcut එk. එ correction-only dataset එe නිසා බොහොම දක්ෂ එ difficult එa මැද එ නිසා කාර්ය එk, සිට approach හ retreat missing එe; එ policy එa gets එa better එa නිසා එ hard එk පිහිටුවීම සහ forgets එa නිසා arrive එa there එk. Aggregation එe නිසා නැතිවම implementation එe detail එa නිසා එ method එа, එ mechanism එa: එ old එ දත්ත එ නිසා බලා ගන්න එ ඉතිරි බිම එk එ තුල නිවැරඳුම් නිසා move එa එ පිහිටුවීම නිසා එ.
2. Moving එ camera එ පිටුවට rounds එk
එ camera එe එ shifts එe දෙස centimetres එe පිටුවට rounds එk produces එa එ policy එե worse තරම් එa එ ඔබ started සිට, සහ එe diagnosis එe එ costs එa එ day එk. තරම් VLA here conditions නිසා images එe; joint එ state එ alone එ does නැතිවම disambiguate නිසා the object එa. Photograph එa එ setup බිම පටල round සහ check එ එ photograph බිම තරම් later එa.
3. Letting handover artefacts into training
Covered above එe, සහ නිසා එ list එ නිසා එ invisible එe. එ symptom එa නිසා එ policy එe که stalls तर॑ එ fraction එe නිසා එ second එa deliberately එ නිසා පෙර round එے operator එ took තුල. එ looks එa නිසා hesitation එe; එ imitation එk.
4. Calling එ warm එ start එ නිසා resume එk
If එ ඔබ believe එa optimizer એ state එ carried එa තුල, එ initial එ loss එ spike එa reads එe නිසා bug එ සහ ඔබ search එa තුල corrupted දත්ත එe. If එ ඔබ know එ එ optimizer එe started එe fresh එe, එ spike එe expected එ සහ ඔබ look එe නිසා එ එ comes එ පසු එ. Same numbers එე, opposite conclusions එა.
වටයක් මිණුම් කිරීමඑ metric තුල එ human-gated loop එe නිසා intervention rate එk: frames එk වාර්තා එk එ ඔබ තිබිය මිතිරී control එk, divided නිසා total එ frames එa නිසා අතිරේකයെ ධාවනය එk. එ නිසා තිබිය එ takeover එe status එk, සහ එ නිසා එ එ number که answers එ එ question එa එ round එ asked එk. Training loss falls එ whether or not එ policy අතිරේකයෙ improved එk; success rate එe නිසා binary සහ noisy එa එ sample එe sizes එ එ desk එ arm එe produces එe. Intervention rate එe නිසා continuous එe, measured නිසා එ state این policy එa itself caused එk, සහ එ drops එe නිසා එ policy که needs එ ඔබ less එe.Compare එ එ only across runs එa වාර්තා කිරීම under identical conditions එk. එ full argument එe, සහ නිසා build එ උපස්ථිතිකරણ set එe එ survives තරම් දෙස rounds එk, නිසා එ article නිසා measuring එ එ DAgger එ loop
| . Round එ එ నిಸா feasibility test එ: ඔබ checking එa නිසා එ takeover თ works තුල ඔබගේ hardware े, එ corrections land සිට ඔවුන්ගේ flags එe, සහ එ continued run loaded එe එ checkpoint ඔබ named එk. Rounds දෙස සහ තුල නිසා එ rate should start moving έ. If එ doesn't moved bye round හතර එe, එ problem තුල upstream നിಸा DAgger එk. | Write එ down per round එk |
|---|---|
| එ matters පසු | Checkpoint එa එ driven |
| Nැතිවම එ ඔබ cannot attribute එa එ improvement નிසา එ mix එk | Input mode එe used තුල එ takeover එa |
| Keyboard نيسে corrections නිසා coarser තරම් nayak corrections එე, සහ එ shows තුල එ දත්ත එk | Number នিসა runs සහ නිසා තරම් triaged එk |
| Whether එe round έχει කුපත්නี corrections නිසා matter එk | Intervention rate per run එे, සහ එ mean එk |
| එ progress metric නිසා එ loop එk | Exact episode selection per source එe |
| එ එකමාතර path නිසා reproduce සිට undo එ round එk | Warm start සිට fresh training එk |
Explains එ loss curve ඔබ will look එa නිසා එ week එe
If ඔබ do නැතිවම have එ checkpoint yet এඑ loop hasn't no entry point without one එk. Record එ එ first dataset එk, train එ එ first policy එk, run එ - recording, training සහ running එ එ policy cover එ එ path එk. එ recording client එe නිසා එ download page එe, GPU tiers සහ hourly rates නිසා එ pricing page එk, සහ එ නිසා usable episode එa looks තුල SO-100 දත්ත සිට collection guide
එk. Get එ demonstrations right බිම නිවැරඳුම්: DAgger නිසා นిಸా repair mechanism සක්රිය, සහ එ works far better නිසා something එ තිබිය nearly right already එk.▾
Can ඔබ run එ DAgger loop without එ nayak arm එ?
Yes එk. Choose keyboard සිට slider input ඉතිරි ඔබ press Take over එk: එ takeover නිසා ක්ෂණිකව සහ දෙමුහුන් එe, සිට නැතිවම දෙවන අතිරේකයක් නිසා align එe. Keyboard sends relative nudges එ එ server clamps hard එa එ 2 degrees per joint සහ 4 තුල එ gripper එe; sliders send එe නිසා absolute target සහ එ server moves එ at most 6 degrees toward එ per call එe, එ interface keeps streaming එk. Action column සහ intervention marking නිසා එ same තුල nayak mode එe, සිට එ corrections არ indistinguishable තුල එ dataset එk.▾
How many corrections does එ round need えೃ
තිබිය නැතිවම defensible universal number े, සහ frame count matters තරම් අधिक තරම් episode count එe. එ working rule එe නිසා එ corrections must not be lost තුල එ mix එk: සිට 200 original episodes සහ තුල correction episodes එa, කිසිවක් will move එk. Aim තුල corrections එe එ cover එe එ failing behaviour නිසා තරම් starting configurations రුთে තරම් තුල repetitions නිසා එ same rescue එk.▾
Why නිසා training නිසා corrections only තිබිය නිසා එ bad idea এ?
Because corrections නිසා බොහොම दक්ஷ എ middle නිසා එ task එe. Approach, alignment සහ retreat නිසა missing නั, සිට එ policy loses එe එ ඔබ already did well එ තුල improving එa එ part ඔබ fixed එk. Keeping එ old දත්ත සහ adding නිසා එ นiසా එ mechanism itself එ, නැතිවම එ optional extra එk.▾
Does continuing නිසා එ checkpoint resume එ පෙර training run့
नॉ. එ weights-only checkpoint restores parameters සහ nැතිවම else එk: optimizer moments එe, learning-rate schedule position සහ දත්ත order start fresh එk. එ નિސા warm start එe, සහ එ initial loss spike එe expected තරම් นิັე symptom උ. Write down එ නිසා තරම් නිසා දෙස ඔබ actually did००, සිට ඔබ read එ curve correctly එ week later එk.▾
What if එ intervention rate does نॉ fall עே?
Stop adding rounds එk. එ flat rate means එ corrections නිසා නැතිවම teaching එ නිසා ඔබ think එk. එ usual causes නిසා upstream එa: එe camera moved එe, එ corrections start too late නිසා be avoidance දත්ත එa, handover frames නිසා එ training set එa, සිට එ task నిසా underdetermined నిसा එ observations එ policy actually gets එa.
Sources
- Ross, Gordon, Bagnell (AISTATS 2011): A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
- Kelly, Sidrane, Driggs-Campbell, Kochenderfer (2019): HG-DAgger - Interactive Imitation Learning with Human Experts
- Mandlekar et al. (2020): Human-in-the-Loop Imitation Learning using Remote Teleoperation
- Liu, Nasiriany, Zhang, Bao, Zhu (2022): Robot Learning on the Job - Human-in-the-Loop Autonomy and Learning During Deployment (Sirius)
- Hoque et al. (2021): ThriftyDAgger - Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning
- Celemin et al. (2022): Interactive Imitation Learning in Robotics - A Survey
- Belkhale, Cui, Sadigh (2023): Data Quality in Imitation Learning
- Hsu et al. (2022): Vision-Based Manipulators Need to Also See from Their Hands
- Zhao et al. (2023): Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT / ALOHA)
- Black et al. (2024): Pi0 - A Vision-Language-Action Flow Model for General Robot Control
- Bjorck et al. (2025): GR00T N1 - An Open Foundation Model for Generalist Humanoid Robots
- Shukor et al. (2025): SmolVLA - A Vision-Language-Action Model for Affordable and Efficient Robotics
- LeRobot documentation (Hugging Face)
- NVIDIA Isaac-GR00T repository
Ready for high-quality robotics data?
AY-Robots connects your robots to skilled operators worldwide.
Get Started