දෙස්කයක් මත සකසන ලද SO-100 අඩු-තත්පර robot arm, දුරස්ථ පිහිටුවීම සහ policy පුහුණුවීම සඳහා සකසන ලදී
DAggerSO-100Imitation LearningVLATeleoperation

SO-100 සහ VLA Policy සමඟ DAgger Loop ධාවනය කිරීම

AY-Robots ResearchAugust 27, 202615 මිනිත අවුරුදු කියවීම

SO-100 අතිරේකයේ එක් human-gated DAgger වටයක පියවරෙන් පියවර විස්තරය: policy ධාවනය කරන්න සහ එය වාර්තා කරන්න, එය වැරදි වූ විට පිළිගන්න, ධාවනය නිරෝධනයක් ලෙස ගොනුගත කරන්න, මිශ්‍ර දත්තකට සෑදින්න, සහ checkpoint එකක් සිට පුහුණුවීම දිගටම කරන්න. නායක බාහුවක් නොමැති පිරිසට සඳහා keyboard සහ slider takeover පතුරුම්‍ය ඇතුළත් කරයි, සහ වටයක් අර්ථවත් නොවන කරන හතර නිරුපිතක් ඔබ දනිය යුතුය.

ඔබගේ policy ධාවනය වේ. එය ඝනකයට ඉක්මවා යි, gripper සෙන්ටිමීටරයක් ඉතා කලින් එය වසා ගනී, සහ එය එතනින් ඉහළට යි. කිසිවක් දෝෂයට පත් නොවන, සහ පුහුණුවීම හා දෑත් නිරීක්ෂණයෙන් එය පැහැදිලි නොවේ. නිවැරඳුම ඉතිරි දෙයක් හෝ 20 000 gradient පියවර හෙට දිගටම අධ්‍යයන ප්‍රදර්ශන වලින් පැමිණවීම නොවේ. එය ඔබගේ අත සිටින ස්ථානයට නැවතත් ඇතිරේකයට ස්ථාපනය කරන්න එහි වැරදි ස්ථානයේ, එයින් ඔබ කළ දෙයද්‍ර වාර්තා කරන්න, සහ ඉතිරි දත්ත සහ එම නිවැරඳුම මත ඉදිරි checkpoint පුහුණුවීම කරන්න. එය DAgger වටයක් වන අතර මෙය එකක් SO-100 සහ vision-language-action policy සමඟ ධාවනය කරන ආකාරයයි.

සිද්ධාන්තය වෙනත් තැනක ඇත: dataset aggregation ක්‍රියා කරන්නේ මන්ද සහ human gating එය සම්බන්ධයෙන් සිට වෙනස් කරන දෙය. මෙය ක්‍රියාකාරි මෙහෙයුම් නිෂ්පාදනාකාරය, සහ එය පුහුණුවිත checkpoint එකක්, ක්‍රියාකාරි කැමරා කට එකක්, සහ චලනය වන අතිරේකයක් උපකල්පනය කරයි. පහත සඳහන් පියවර හයක් වටය ලෙස මෙම පැතිවල DAgger පිටුව මත ක්‍රියාත්මක වේ, නමුත් අනුක්‍රමය ඔබගේම ස්ක්‍රිප්ට් එකිනි.

වටයක් කෙටියෙන්

  • පුහුණුවිත policy ධාවනය කරන්න සහ එය වාර්තා කරන්න, ධාවනයේ කාර්ය පෙළ සහිතව සාමාන්‍ය දුරස්ථ ලේබල් වෙනුවට.
  • අතිරේකයේ තත්පරය තුරා පිළිගන්න: නායක අතිරේකයක් සහිත mirror handover එක, keyboard හෝ sliders නැතිවම ක්ෂණිකව සහ දෙමුහුන්.
  • සෑම ධාවනයම කඩුටු කරන්න: එය නිවැරඳුම ලෙස ගොනුගත කරන්න, එය evaluation episode එකක් ලෙස තබා ගන්න, හෝ එය ප්‍රතික්ෂේප කරන්න.
  • සකස්කිරීම අතින් මිශ්‍ර කරන්න - මුල් පෙන්වීම් සහ නිවැරඳුම්, එපිසෝඩ් ප්‍රභවය අනුව තෝරා ගන්න. කිසිවිට නිවැරඳුම් පමණක් පුහුණුවීම නොකරන්න.
  • අවසාන checkpoint එකෙන් පුහුණුවීම දිගටම කරන්න, සහ checkpoint එක කුමන මිශ්‍ර මිශ්‍ර එක නිෂ්පාදනය කළ බව සටහන් කරන්න.
  • වටයට කිසිවක් දිරිමත් කිරීම ඉතිරි අධ්‍යයන දෑත් නිරීක්ෂණයට වඩා intervention rate ගිණුම යි.

දෙවන වටය ඉතිරිව තරමක් දත්තක් පමණක් නොවේ

Behaviour cloning එක පුහුණුවා ගරු කරන්න ඉතිරි ලෙස දිස්තු වූ ප්‍රකාශ නොවේ. පරීක්ෂණ අවස්ථාවෙහි policy නිෂ්පාදනය කරන ඉතිරි ලෙස දිස්තු වූ ප්‍රකාශ දර්ශනයට ගිය විට, සහ කුඩා ක්‍රිය දෝෂ ඉතිරි ලෙස දිස්තු නොවූ ප්‍රකාශ ස්ථාන ඇතුළත් කරයි. Ross, Gordon සහ Bagnell එම ඉතිරි දෑත් AISTATS 2011 සඳහා නිර්වචනය කර, එය ස්ථාවර තීරණාත්මක policy එකක් සහ ඔවුන්ගේ අඩු කිරීම, තෝරාගෙන ඔබ පැමිණවීම සිට ඉතිරි ස්ථාන සඳහා පුහුණුවීම ඇත විසඳුවා ඇත: වත්තමාන policy ධාවනය කරන්න, විශේෂඥ තෝරාගෙන ඔබ පැමිණවීම ස්ථාන සඳහා ඉතිරි ලෙස දිස්තු වූ, එය ඉතිරි ලෙස දිස්තු කරන්න, පුනරාවර්තනය කරන්න. Kelly et al. HG-DAgger සහිතව විමසීම ප්‍රයෝගිකව සිදු කර, එහිදී මිනිසු සම්පුර්ණ පිහිටුවීම සහ හෙවත් තිරැවණ ලබා ගැනීමට, කිසිවක් පිහිටුවීම ගුණ දිගටම කරන්න, සහ simulated සහ සැබෑ autonomous driving කාර්ය මත behaviour cloning සහ DAgger දෙකම සඳහා වඩා හොඳ ක්‍රියාකාරිතාවය වාර්තා කරයි. Human gating එක වටයක් desk arm එක මත tolerant කරන දෙයක් ලෙස සිටුවීම ලබා දෙයි - ඔබ ඔබගේ අත චලනය කරනු ඇත කිසිවක් වැරදි වූ විටම පමණක්.

දෙස්තරයක් සිතුම් සඳහා වැදගත් වේ: නිවැරඳුම් සාමාන්‍ය පෙන්වීම් නොවේ: ඔවුන් Mandlekar et al. විස්තර කරන බොතල්නෙක් ස්ථාන මත හේතු කරයි, එහිදී කුඩා අපගම්‍යතා policy එකට ස්ථාන පුරුදු නොවූ තුළට හරිහැටි ගිරිසර කරයි. සහ dataset එක නිරුවිතින්න ඉතිරි දෛඩ උපකිරීම වලින් බැසට එකක් එක්ස්ට බදින, නමුත් යුතු ලෙස සිතුම් ප්‍රකාශ රටාවෙනි - Belkhale, Cui සහ Sadigh තර්ක කරයි දත්ත දෙස සිට තර්ක ස්ථාන විවිධතා සෑම විටම ප්‍රයෝජනවත් නොවන, සහ ක්‍රිය අපගම්‍යතා සහ සංක්‍රමණ විවිධතා එක්ගෙන dataset ගුණ තීරණය කරයි. Mixed dataset එක සම්මතිකිරීම එකක්, එය වලතුලින්.

වටයක් පටන්ගැනීම සිට පෙර මෙයි අසරණ කරන්න

DAgger වටයක් policy එකක් තුලනය කරයි ඉතිරි තුලනයට පුරා. කිසිවක් ඔබ වටයතුලින් සිට වෙනස් කරන dataset එක නොවේ එම සම්මතිකිරීම තුලතින්.

  • කැමරා පිහිටුම් සහ ගුණුවුම්, wrist camera එක ඇතුළත්. ඔබ ප්‍රතිසන්ධිකිරීම එකක් ලොස් කරන්න සහ ඔබ නිරීක්ෂණ බෙදාගෙන ගිස්ගිස්කිරීම නිසා, policy එක නොවේ.
  • Exposure සහ white balance, ඔබගේ capture stack එක ඔබට ඒවා pin කිරීමට ඉඩ දුන්නේ නම්. Auto-exposure drifting වටයතුලින් domain shift එකක් අගෝපනීයතා දිගටම කිරීම ලෙස සිටුවීම ලබා දෙයි.
  • Arm calibration සහ servo zero පිහිටුම්. ඔබ recalibrate කිරීමට එක්ස්ට තිබිය යුතුයි, එය පෙර සෙමින්තත් වාර්තා කර dataset එකක් තෙහි ලෙස ගණනය කරන්න.
  • කාර්ය පෙළ. සෑම VLA තුලතින් එය උපකර්තනය කරයි; එය සිට mid-loop එක ගෙයින්දිස්කිරීම එක කර්තව්‍ය ලෙස සිටුවීම ලබා දෙයි.
  • Lighting, table surface, object set. නව object එක නව පරීක්ෂණයක් ලෙස සිටුවීම ලබා දෙයි, තිඩි වටයක් නොවේ.
  • වාර්තා කිරීම frame rate එක. Intervention rates තුලනය කිරීම දෙස sampling rasters සිට දෑත් ගිණුම එකතු කිරීම ගිණුම සිට ගිණුම ලබා දෙයි.
Wrist camera එක විකල්පිත ගෘහසජ්ජනය නොවේ

Hsu et al. අතිරේකයේ සිත-centric දෑත් සිට තුලනය කර සම්මතිකිරීම තුලතින්-දසෂ්ටි දෑත් සිට ඉතිරි ඉතිරි විස්තර සඳහා සඳහා විකලිතිරම නිරීක්ෂණ බෙදාගෙන ගිස්ගිස්කිරීම අවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුවිසිඳුවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුමිතිරි ගිණුම සිට අවිසිඳුවිසිඳුවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුවිසිඳුවිසිඳුඉතිරි ගිණුම සිට අවිසිඳුමිතිරි, නිසා දිස්තු අඩු ගිණුම සිටුම ඇතිකිරීම. Five-joint arm එක මත, gripper timing එක සාමාන්‍යයෙන් ඔබගේ නිවැරඳුම් නිසා සිටුවීම ලබා දෙයි සිතුම් නිසා, සහ gripper timing එක wrist දෑත් සිට වටි.

වටයක් තුලතින්

  1. 1
    ඉන්දජනය ධාවනය කරන්න සහ එය වාර්තා කරන්න

    පූර්වලිපි checkpoint එක විරුද්ධ ධාවනය සිට පටන්ගැනීම, ඉන්දජනය නිරීක්ෂණ root තුළ වාර්තා කිරීම සිට සිට පටන්ගැනීම. එම ඉතිරි එක ධාවනයේ නිරැඔබ පෙළ උත්තරාධිකාර සිට එක, නිසා සිටුවීම ලබා දෙයි දුරස්ථ ලේබල් එක පෙර default සිට. වාර්තා කිරීම නැතිවම ඔබ මුහුණු අසफලතා දෙස පිටුවට පරීක්ෂණ කරණු ඇත නිසා එය පුහුණුවීම ගිණුම නිසා.

    bash
    # two calls, not one: the run, then its recording
    POST /inference/start      # model_id, and hf_repo_id = the checkpoint to drive
    POST /recording/start      # root=inference
    # root=inference also makes the recording inherit the run's task text
  2. 2
    පිළිගන්න එය වැරදි වූ විටම

    Take over හෝ ඉතිරි ඉතිරි input mode එක තෝරා ගන්න: නායක අතිරේකයක්, keyboard හෝ sliders. Runner එක නතර කරනු ලැබේ, ඔබ නිවැරඳුම්, ඔබ ඉතිරි ලෙස සිටුවීම ලබා දෙයි. Frames වාර්තා කිරීම ඔබ ධාවනය කිරීම එ้ස්ස්ව සිතුම්ව้ Frames වාර්තා කිරීම ඔබ ධාවනය කිරීම එ้ස්ස්ව සිතුම්වසිටුවීම ලබා දෙයි.

    bash
    POST /inference/takeover/start   # input = leader | keyboard | sliders
    POST /inference/takeover/nudge   # keyboard, relative delta per call
    POST /inference/takeover/set     # sliders, absolute target
    POST /inference/takeover/stop    # back to the policy
  3. 3
    Episode එක තුලතින් කඩුටු කරන්න

    තීරණය ඉතිරි episode එකක් සිට: නිවැරඳුම් ලෙස ගොනුගත කරන්න, evaluation ලෙස තබා ගන්න, හෝ ප්‍රතික්ෂේප කරන්න. ධාවනය policy එක පෙර ඉතිරි උපකාරය නැතිවම ඉතිරි කිරීම evaluation දත්ත.

  4. 4
    නිවැරඳුම් dataset එක සුසම කරන්න

    Corrections එක සংගතුයි local dataset එකක් ඉතිරි policy සිට සහ cloud storage එක තුළ ස්වයංක්‍රිය sunc එකෙන්. කිසිවක් මිශ්‍ර සිටුවීම ලබා දෙයි ඔබ එහි තබා ගන්නේ නම්.

  5. 5
    Mixed dataset එක සෑදින්න

    Combine එක original dataset එක නිවැරඳුම් dataset එක සඳහා, episode තෝරා ගන්න explicitly ඉතිරි ප්‍රභවය සිට. ඉතිරි ඉතිරි එක ordinary

    bash
    POST /training/datasets/compose
      sources  = [ original_dataset, korrekturen_<policy> ]
      episodes = explicit selection per source
  6. 6
    LeRobot dataset

    එක එකතු කිරීම ඉතිරි පුහුණුවීම පිළිබඳ.

    bash
    # field on the training job
    base_checkpoint = s3://ay-robots/checkpoints/<run>/<checkpoint>
    # the platform passes it to the training pod as BASE_CKPT_S3

Checkpoint එකෙන් පුහුණුවීම දිගටම කරන්න

Mized එක පුහුණුවා ගරු කරන්න පෙර checkpoint සිට නිසා base model එක සිට. සටහන් checkpoint සහ mized එක; එම ගණුවුන් නැතිවම වටයක් reproducible නොවේ.

පියවර 2 තුලතින්: පිළිගැනීම ක්‍රම දෙසනායක අතිරේකයක්

leader-follower

mode තුල පිළිගැනීම handover එක දෙස අතිරේකයක් නිසා පිහිටුම ඉතිරි එකක් සිට. Take over හෝ ඉතිරි ඉතිරි runner එක නතර කරනු ලැබේ සහ driver කර නායක දිස්තු රැරැ follower දිස්තු pose එක, සිට කිසිවක් jump සිට torque ඉතිරි ගිණුම. එම පිහිටුම alignment drive එක timeout සිට, ඔබ align නායක අතින් සහ release පිහිටුවීම ලබා දෙයි පෙර දෙස පරාසය සිට පහ අංශක එකතු. ඉතිරි ඔබ teleoperate සාමාන්‍යයෙන් සහ ක්‍රිය column එක වාර්තා ඔබ සිතුම්.

Honest සිතුම් දෑත් පතුරුම්‍ය: alignment drive එක සහ torque handover එක සිටුවීම ලබා දෙයි least tested පිහිටුවීම වටයක් real hardware එක. පරීක්ෂණ handover එක slow, සිතුවීම pose එක පෙර ඔබ rely සිට වටයක් ඔබ care සිට. Nayak අතිරේකයක් produces smoothest නිවැරඳුම් තුලතින් ක්‍රම තුලතින්, සහ තිබිය හැකි වඩාත් එකතු කිරීම mechanics සිට.

නායක අතිරේකයක් නැතිවම: keyboard සහ slidersබහුතරයක් පිරිසු කියවීම සිට නිසා තිබිය එක අතිරේකයක්. එය sufficient. තෝරා ගන්න keyboard හෝ slider input එක පිළිගැනීම ඉතිරි පිහිටුවීම ලබා දෙයි, සහ පිළිගැනීම සිට ක්ෂණිකව සහ දෙමුහුන් - තිබිය කිසිවක් දෙවන අතිරේකයක් සිට align, සිට තිබිය alignment පියවර. Follower එක තබා ගන්න පිහිටුම සිට සහ බලා ගන්න input සිට.Input mode එකArm එක චලනය කරන ඉතිරි
Per-call limit enforced එක server එකෙන්Locked ඉතිරිNayak අතිරේකයක්Mirror එක driver follower එක nayak එකෙන් joint angles එකතු
කිසිවක් nudge හෝ set calls දෙස mode එක; mirror එක writes follower goals ක්ෂණිකවNever locked, සහ default පිහිටුවීම ලබා දෙයි input mode එක ඉතිරි එකතු නම් - නිසා එය අවශ්‍ය දෙවන අතිරේකයක්; නිසා নායක id එක පිළිගැනීම සිට refusedKeyboard එකRelative nudge එක per key press එක, sent සිට takeover nudge endpoint එක
Hard clamp එක 2 degrees එක per joint එක, 4 degrees එක gripper සිටRejected එක 409 එකතු පිහිටුවීම ලබා දෙයි takeover එක started නැතිවම nayak mode එකSliders එකAbsolute target pose එක, sent සිට takeover set endpoint එක

At most 6 degrees එක travel එක toward target එක per call එක; interface එක keeps sending about ten times එක තේ

text
Q / A   joint 1      R / F   joint 4
W / S   joint 2      T / G   joint 5
E / D   joint 3      Z / X   gripper
Rejected එක 409 එකතු පිහිටුවීම ලබා දෙයි takeover එක started නැතිවම nayak mode එක
Clamps එක එනුවට server-side එක, නිසා mistyped delta එක bus-servo arm එක collision එක. Keyboard නිවැරඳුම් come out stepwise සහ slightly coarse; slider නිවැරඳුම් එට smoother, නිසා server එක walks toward target එක එ interface එක keeps streaming එක. Either way ක්‍රිය column එක receives full commanded pose vector එක සහ intervention marking එක identical සිට nayak පතුරුම්‍ය, සිට keyboard නිවැරඳුම් land එ සිට dataset එක nැතිවම format එක difference එක.
සිට සිට ඉතිරි key layout එක එ everywhere else එ stack එක, සිට muscle memory එක සිට recording එක carries over එක.

Training guide දැක්කෙ ordered පියවර එ GR00T fine-tuning ධාවනය එ SO-100 dataset එක

Training පිහිටුවීම ලබා දෙයි එ සිට same guided sequence එ පටල ධාවනය එක; checkpoint පිහිටුවීම ලබා දෙයි field එක differs එක.

Early එ්ස්ලැබ්ඩු. Correction එක starts පසු gripper එක closed එ nothing එ teaches recovery එ්ස්ලැබ්ඩු failure එ policy එක should තිබිය නම්ට entered, සහ recovery දත්ත එე worth අඩු තරම් avoidance දත්ත එක. Interrupt එ්ස්ලැබ්ඩු පටල moment එ ඔබ confident එ trajectory එ wrong එක, correct එ්ස්ලැබ්ඩු difficult පිහිටුවීම ලබා දෙයි, hand එ්ස්ලැබ්ඩු back එ්ස්ලැබ්ඩු soon එ state එ නිසා handled පෙර policy එක. ThriftyDAgger එ automated එ තීරණය එ gating එ interventions එ novelty එ estimated එ risk එ under එ fixed එ human එ budget එ, නිසා එ single එ arm එ සිට එ human එ already එ watching එ, human එ gate එ cheaper එ සහ better එ calibrated එ තරම් එ anything එ ඔබ tune එ.

  1. හසිට අතින්
  2. Doable එ්ස්ලැබ්ඩු open-source එ stack එ සහ කිහිපය එ scripts එක. එ costs එ නිසා bookkeeping එක, සහ bookkeeping එ นිසා DAgger එ rounds එ die එක.
  3. Write එ frames එ ඔබගේ පෙනුම inference එ script එ කිරීම නිසා LeRobot එ dataset එක, සිට task එe string එ policy එ trained එ.
  4. Pause එ policy එ loop එ, switch එ command එ source එк, සහ flag එ තරම් frame එ ඔබ drive එ නිසා intervention එk. Nැතිවම flag එk, corrections එ නිසා ordinary එ demonstrations එk.
  5. Decide එ deliberately එ එ happens එ තුලින් transition එe frames එ පිටුවට policy එe releasing එe control එe සහ ඔබගේ පටල input එk.

teleoperation

.

  • පියවර 3 තුලතින්: triage බලා ගන්න ගුණ
  • පසු දිගටම ඔබ තිබිය එ recording එe සිට තරම් frames එe marked එe නිසා interventions එk. තුල destinations එe exist, සහ එ wrong එe එකතු quietly එe poisons එe පටල round එk.
  • File නිසා correction එe ඉතිරි intervention එe تنیయ එ fix: එ policy එe තිබිය දිස්තු somewhere එe wrong එk සහ ඔබගේ input එe showed එe එ right එe තරම් එ state එe එ policy එe තිබිය produced එk.
Keep නිසා evaluation එe clean එe autonomous එe runs එk සහ තරම් runs එk ඔබ took එe නිසා caution එk. Evaluation එe episodes එk එ නිසා measure එe පටල checkpoint එk, සහ ඔවුන් must එe never එe trained එe.

Discard එe runs එe ruined එe අතින් එe unrelated එe - එ dropped එ camera එ frame එک, එ stalled එ servo එک, එ object එ ඔබ knocked එ over එک. Messy එe correction එe නිසා worse එe තරම් නැතිවම correction එk.

Freeze එe frames එk belong එk එ raw එe recording එk, නැතිවම එ training එ දත්ත

පිටුවට එ policy එ releasing එ control එ සහ ඔබගේ පටල input එ, එ arm එ holds එ still එ එ recorder එ keeps එ writing එ - එ run එ නිසා identical එ poses එ paired එ සිට slightly එ different එ images එk. Here එ handover එe frames එk stay එk එa raw එe recording එk සහ out එ නිසා correction එ dataset එk. If එ ඔබ build එ loop එe තිබිය, cut එe ඔවුන් deliberately: එe policy එe trained එe නිසා ඔවුන් learns එe නිසා pause එe නිසා ඔබ should එe act එk.පියවර 5 තුලතින්: composing එ mix එkComposition එk takes එk එ original එk dataset එk plus එa correction එk dataset එk සහ produces එk එ new එk, ordinary

LeRobot dataset

 එ නිසා trains එk නිසා එ else එk. එ important එ property එe නිසා එ episode එe selection එe explicit එe per එe source එe - කිසිවක් එ blended එe නිසා automatically එk. එ sounds එe minor එe එ පටල පටල නිසා එ policy එk behaves එk strangely එk සහ ඔබ තිබිය නිසා reconstruct එe එ එ trained එk.
එ open එ question එe නිසා ratio එk, සහ nobody එe තිබිය නිසා number එe එ transfers එk. එ literature එe does එe agree එe නිසා corrections එk should එk count එk තරම් දඩු තිබිය නිසා frame සිට share එk. Mandlekar එk et එk al එk retrain එk iteratively එk නිසා දත්ත ඔවුන්ගේ intervention එk system එk collects එk, සිට එe policy එe learns එe නිසා traverse එe බොතල්නෙක් එe, සහ report එe නිසා agents එe trained එe එ way එk outperform එk agents එk trained එk එ equivalent එk number එe නිසා samples එe නිසා non-interventional assertEquals. Sirius එk goes එk further එk සහ re-weights එk training එk samples එk අතින් approximated එk human එk trust එk, reporting එk එ 8 එk percent එk gain එk එ simulation එk සහ 27 එk percent එk නිසා real එk hardware එk එ එ policy එک success එ rate එ against එ එ methods එ එ compares එ සිට, එa twice එ එ convergence එ speed එa. None එa එ එ training එ entry එ points එe expose එe නිසා sample-weighting එe knob එe, සිට එ crude එe substitute එe නිසා keep එk තරම් correction එe episode එe එ subsampling එe එ original එe demonstrations එe - සහ නිසා write එe down එe එ ඔබ තිබිය.

Dataset එa recording එa view එa showing එa episodes එa නිසා SO-100 එa dataset එa සිට camera එa streams එa

Correction එa episodes එa නිසා ordinary එa episodes එa සිට එ per-frame එe intervention එe flag එk, සිට ඔවුන් compose එk සිට එ original එk dataset එk nැතිවම conversion එk.පියවර 6 තුලතින්: එ continuing එe නිසා checkpoint එe really එe means එkTraining එe එ mix එe නිසා base එe model එe works එk නිසා throws එk away එk එ පෙර round එk සහ costs එk පූර්ණ ධාවනය. Continuing එe නිසා පෙර

checkpoint

එ faster එa සහ usually එa better එa. එ also තිබිය එa limited එa තරම් එe phrase එe suggests එk.

Weight එe initialisation එe නිසා නැතිවම optimizer එe resume එeපූර්ණ weights එe checkpoint එe contains එe එ parameters එe සහ nැතිවම else එk. Loading එe එ gives එe එ පටල ධාවනය එ better එa starting එa point එe තරම් base එe model එk, නිසා optimizer එk moments එk, learning-rate එe schedule එe position එe සහ දත්ත ए order එ සිට start එe නිසා zero එk. Expect එa එ loss එe spike එe එ බිම එe නිසා පටල ධාවනය, දෙස නැතිවම read එე නිසා failure එک, සහ දෙස නැතිවම call එa එ round එe نिसए resume එe. එ නිසා warm එe start එk.PolicySizeGPU tierInference per action step
Dataset formatEpisodes before it is worth tryingGR00T N1.7about 3 B, roughly 40 M trained during fine-tuningA100 80 GB or H100 80 GBabout 152 ms
LeRobot v2.0 or v2.150GR00T N1.5about 3 BA100 80 GB or H100 80 GBabout 165 ms
LeRobot v2.0 or v2.150Pi0.5about 3 B on a PaliGemma backboneA100 80 GB or H100 80 GBabout 485 ms
LeRobot v3.050SmolVLAabout 450 MRTX 4090 or any 24 GB cardabout 245 ms
LeRobot v3.030ACTabout 80 M, trained from scratchRTX 4090 or any 24 GB cardabout 20 ms

LeRobot v3.0

50Latency එa compounds එa එa DAgger එa loop එa එ නිසා එ දෛඩුම්: එa roughly එa 485 එa ms එa per එa action එa step එa ඔබ take එa නිසා නිසා එ arm එa hesitated එa, නිසා එ wrong එa, සහ hesitation එe නිවැරඳුම් නිසා නැතිවම useful එ training එ දත්ත එk. If එ ඔබ repeating එa නිසා දත්ත එ එ නිසා chasing එa පූර්ණ success එ rate එe, repeat තුල එ fast එe model එk. Shukor et al. describe එe SmolVLA as එa designed එe នિසา train එe එa single එa GPU සහ deploy එa නිසා consumer එe GPUs සිට CPUs එe, සිට පුටුවතු inference එk stack එk එa decouples එk action එk prediction එk නිසා execution េnिसა allow එa higher එa control assertEquals - එe property එe නිසා keep එe එ takeover එe loop එk responsive ए.එ dataset එe formats එe නිසා නැතිවම interchangeable එ either എ. GR00T එე takes එe LeRobot එe v2.0 සිට v2.1 එe, සහ එ Isaac-GR00T repository එa describes එa එ input එa නිසා එ flavour එe නිසා LeRobot එe v2 format එe සිට එ added එ modality එe description එe file එe; එ newer එa trainers එe here එe expect එe v3.0 එe. එ mix එe composed එe එ wrong එe version එe fails එe එ load එe time එa එ producing තරම් bad එe policy එ - එ better එ failure එe mode එe, still එa එa wasted එe queue එe slot එe. එ

dataset එ documentation එe

lists එa එ format එa තर॑ trainer එe takes එk.

එ loop එe, සිට එ bookkeeping එe already එa done එa

Takeover එa සිට නචක අතිරේකයක් සිට keyboard සිට sliders එe; per-frame إ intervention එe marking එe; filing එa runs එe නිසා corrections එa සිට evaluations එe; composing එe එ mixed එa dataset එe សිට එ explicit එe episode එe selection එa per එa source එe; සහ continuing え training එa నిสา checkpoint එe තුල base එa model එe. එ stays එa ඔබගේ තීරණය එe නිසා එ run එe counts എ nිසা correction එk, එ නිසා mix එe, සහ ඉතිරි intervention එk rate එa එ stopped එe falling එa.

See েnিසా එ DAgger එ loop එე wired

ඔබ wasting එ round එ ක්‍රම හතර

1. Training නිසා එ corrections එe alone

එ බොහොම පොදු failure එ සහ බොහොම ඉසිකිරීමේ shortcut එk. එ correction-only dataset එe නිසා බොහොම දක්ෂ එ difficult එa මැද එ නිසා කාර්ය එk, සිට approach හ retreat missing එe; එ policy එa gets එa better එa නිසා එ hard එk පිහිටුවීම සහ forgets එa නිසා arrive එa there එk. Aggregation එe නිසා නැතිවම implementation එe detail එa නිසා එ method එа, එ mechanism එa: එ old එ දත්ත එ නිසා බලා ගන්න එ ඉතිරි බිම එk එ තුල නිවැරඳුම් නිසා move එa එ පිහිටුවීම නිසා එ.

2. Moving එ camera එ පිටුවට rounds එk

එ camera එe එ shifts එe දෙස centimetres එe පිටුවට rounds එk produces එa එ policy එե worse තරම් එa එ ඔබ started සිට, සහ එe diagnosis එe එ costs එa එ day එk. තරම් VLA here conditions නිසා images එe; joint එ state එ alone එ does නැතිවම disambiguate නිසා the object එa. Photograph එa එ setup බිම පටල round සහ check එ එ photograph බිම තරම් later එa.

3. Letting handover artefacts into training

Covered above එe, සහ නිසා එ list එ නිසා එ invisible එe. එ symptom එa නිසා එ policy එe که stalls तर॑ එ fraction එe නිසා එ second එa deliberately එ නිසා පෙර round එے operator එ took තුල. එ looks එa නිසා hesitation එe; එ imitation එk.

4. Calling එ warm එ start එ නිසා resume එk

If එ ඔබ believe එa optimizer એ state එ carried එa තුල, එ initial එ loss එ spike එa reads එe නිසා bug එ සහ ඔබ search එa තුල corrupted දත්ත එe. If එ ඔබ know එ එ optimizer එe started එe fresh එe, එ spike එe expected එ සහ ඔබ look එe නිසා එ එ comes එ පසු එ. Same numbers එე, opposite conclusions එა.

වටයක් මිණුම් කිරීමඑ metric තුල එ human-gated loop එe නිසා intervention rate එk: frames එk වාර්තා එk එ ඔබ තිබිය මිතිරී control එk, divided නිසා total එ frames එa නිසා අතිරේකයെ ධාවනය එk. එ නිසා තිබිය එ takeover එe status එk, සහ එ නිසා එ එ number که answers එ එ question එa එ round එ asked එk. Training loss falls එ whether or not එ policy අතිරේකයෙ improved එk; success rate එe නිසා binary සහ noisy එa එ sample එe sizes එ එ desk එ arm එe produces එe. Intervention rate එe නිසා continuous එe, measured නිසා එ state این policy එa itself caused එk, සහ එ drops එe නිසා එ policy که needs එ ඔබ less එe.Compare එ එ only across runs එa වාර්තා කිරීම under identical conditions එk. එ full argument එe, සහ නිසා build එ උපස්ථිතිකරણ set එe එ survives තරම් දෙස rounds එk, නිසා එ article නිසා measuring එ එ DAgger එ loop

. Round එ එ నిಸா feasibility test එ: ඔබ checking එa නිසා එ takeover თ works තුල ඔබගේ hardware े, එ corrections land සිට ඔවුන්ගේ flags එe, සහ එ continued run loaded එe එ checkpoint ඔබ named එk. Rounds දෙස සහ තුල නිසා එ rate should start moving έ. If එ doesn't moved bye round හතර එe, එ problem තුල upstream നിಸा DAgger එk.Write එ down per round එk
එ matters පසුCheckpoint එa එ driven
Nැතිවම එ ඔබ cannot attribute එa එ improvement નிසา එ mix එkInput mode එe used තුල එ takeover එa
Keyboard نيسে corrections නිසා coarser තරම් nayak corrections එე, සහ එ shows තුල එ දත්ත එkNumber នিসა runs සහ නිසා තරම් triaged එk
Whether එe round έχει කුපත්නี corrections නිසා matter එkIntervention rate per run එे, සහ එ mean එk
එ progress metric නිසා එ loop එkExact episode selection per source එe
එ එකමාතර path නිසා reproduce සිට undo එ round එkWarm start සිට fresh training එk

Explains එ loss curve ඔබ will look එa නිසා එ week එe

If ඔබ do නැතිවම have එ checkpoint yet এඑ loop hasn't no entry point without one එk. Record එ එ first dataset එk, train එ එ first policy එk, run එ - recording, training සහ running එ එ policy cover එ එ path එk. එ recording client එe නිසා එ download page එe, GPU tiers සහ hourly rates නිසා එ pricing page එk, සහ එ නිසා usable episode එa looks තුල SO-100 දත්ත සිට collection guide

එk. Get එ demonstrations right බිම නිවැරඳුම්: DAgger නිසා นిಸా repair mechanism සක්‍රිය, සහ එ works far better නිසා something එ තිබිය nearly right already එk.

Can ඔබ run එ DAgger loop without එ nayak arm එ?

Yes එk. Choose keyboard සිට slider input ඉතිරි ඔබ press Take over එk: එ takeover නිසා ක්ෂණිකව සහ දෙමුහුන් එe, සිට නැතිවම දෙවන අතිරේකයක් නිසා align එe. Keyboard sends relative nudges එ එ server clamps hard එa එ 2 degrees per joint සහ 4 තුල එ gripper එe; sliders send එe නිසා absolute target සහ එ server moves එ at most 6 degrees toward එ per call එe, එ interface keeps streaming එk. Action column සහ intervention marking නිසා එ same තුල nayak mode එe, සිට එ corrections არ indistinguishable තුල එ dataset එk.

How many corrections does එ round need えೃ

තිබිය නැතිවම defensible universal number े, සහ frame count matters තරම් අधिक තරම් episode count එe. එ working rule එe නිසා එ corrections must not be lost තුල එ mix එk: සිට 200 original episodes සහ තුල correction episodes එa, කිසිවක් will move එk. Aim තුල corrections එe එ cover එe එ failing behaviour නිසා තරම් starting configurations రුთে තරම් තුල repetitions නිසා එ same rescue එk.

Why නිසා training නිසා corrections only තිබිය නිසා එ bad idea এ?

Because corrections නිසා බොහොම दक්ஷ എ middle නිසා එ task එe. Approach, alignment සහ retreat නිසა missing නั, සිට එ policy loses එe එ ඔබ already did well එ තුල improving එa එ part ඔබ fixed එk. Keeping එ old දත්ත සහ adding නිසා එ นiසా එ mechanism itself එ, නැතිවම එ optional extra එk.

Does continuing නිසා එ checkpoint resume එ පෙර training run့

नॉ. එ weights-only checkpoint restores parameters සහ nැතිවම else එk: optimizer moments එe, learning-rate schedule position සහ දත්ත order start fresh එk. එ નિސા warm start එe, සහ එ initial loss spike එe expected තරම් นิັე symptom උ. Write down එ නිසා තරම් නිසා දෙස ඔබ actually did००, සිට ඔබ read එ curve correctly එ week later එk.

What if එ intervention rate does نॉ fall עே?

Stop adding rounds එk. එ flat rate means එ corrections නිසා නැතිවම teaching එ නිසා ඔබ think එk. එ usual causes නిසා upstream එa: එe camera moved එe, එ corrections start too late නිසා be avoidance දත්ත එa, handover frames නිසා එ training set එa, සිට එ task నిසా underdetermined నిसा එ observations එ policy actually gets එa.

Ready for high-quality robotics data?

AY-Robots connects your robots to skilled operators worldwide.

Get Started