
Intervention rate - intervention frames ஐ run frames மூலம் வகுப்பதால் - மனிதனால் கட்டுப்படுத்தப்படும் DAgger loops மிக விலையுயர்ந்த நேர்மையான முன்னேற்ற சமிக்ஞை. இந்த கட்டுரை அதை வரையறுக்கிறது, இரண்டு மனிதர்களும் அதே மதிப்பை கணக்கிட பயன்படுத்தப்படுகிறது, அதை எவ்வாறு பதிவு செய்வது என்பதைக் காட்டுகிறது, மற்றும் நான்கு வழிகளை வேலை செய்கிறது அது உங்களை வழிநடத்துவது: operator habituation, ஒரு drifting evaluation setup, corrections மீது மட்டும் பயிற்சி, மற்றும் ஒரு மனிதன் செவ்வாய்கிழமையிலும் திங்கட்கிழமையிலும் வேறுபட்ட முறையில் திருத்துகிறான்.
DAgger round உணர்வு உபயுக்த நீங்கள் அதில் இருக்கும் போது. நீங்கள் policy ஐ இயக்குகிறீர்கள், grasp ஐ தவறவிட்டபோது எடுத்துக் கொள்ளுங்கள், corrections ஐ சேமிக்கவும், retrain செய்யவும், அடுத்த run - நீங்கள் சொல்வதுண்டு - கொஞ்சம் மெருமெரு என்று. இரண்டு rounds பின்னர் எதுவும் மாற்றியுள்ளதா என்று நீங்கள் கூற முடியாது, ஏனெனில் "கொஞ்சம் மெருமெரு" ஒரு அளவு அல்ல.
Loop ஒரு round க்கு ஒரு எண்ணுக்கு தேவை, மற்றும் ஒரு human-gated loop இல் அந்த எண் கிட்டத்தட்ட இலவசம்: run இன் பகுதி நீங்கள் policy க்கு பதிலாக control இருந்தீர்கள். இந்த கட்டுரை அதை வரையறுக்கிறது, இரண்டு மனிதர்களும் அதே மதிப்பை கணக்கிட பயன்படுத்தப்படுகிறது, அதை பதிவுசெய்யுங்கள், விளைந்த curve ஐ படிக்கவும், மற்றும் நான்கு வழிகளை பெயரிடுங்கள் இது உங்களை வழிநடத்துகிறது எப்போது இது தனிப்பட்டதாக நிற்கிறது. இந்த பகுதி கருதும் கோட்பாடுக்கு, பார்க்கவும் the DAgger explainer மற்றும் the human-gated variant; hands-on counterpart ஆக running a DAgger loop on an SO-100.
Short version
- •Intervention rate = intervention frames / run frames. Frames, episodes அல்ல, மற்றும் denominator நீங்கள் correcting செலவழிச்ச frames உள்ளடக்கியது.
- •மூன்று rounds க்கு ஒரு flat curve அர்த்தம் rounds எதுவும் வாங்கவில்லை - mixture, checkpoint அல்லது task ஐ நான்காவது collect செய்வதற்கு பதிலாக மாற்றவும்.
- •Intervention rate falling ஆக இருப்பது rising success rate அல்ல. Stepping in க்கான உங்கள் threshold drift downward என்று நம்பிக்கை வளரும் போது.
- •Frozen evaluation protocol இல்லாமல் - same start poses, objects, cameras, trial count - நீங்கள் room ஐ அளவிடுகிறீர்கள்.
- •Corrections மீது மட்டும் training state distribution ஐ skew செய்கிறது. IWR இன் பதில்: intervention மற்றும் non-intervention samples சமம் விகிதம் draw செய்யவும்.
ஏன் round க்கு ஒரு எண் தேவை
DAgger உள்ளது distribution problem காரணத்தால், மற்றும் distribution problems கண்ணுக்கு கணக்கிடப்பட முடியாது. Ross மற்றும் Bagnell காட்டினர் imitation learning fixed demonstration set மீது leading compounding errors மற்றும் regret bound growing quadratically time horizon இல்: demonstrator visited states கள் கூட learner கூட என்பது ஒரு சமயம், data இல் எதுவும் மீண்டும் வர எவ்வாறு பற்றி சொல்லுகிறது. DAgger fixes இதை iterating மூலம் - run the policy, label the states it visits, aggregate, retrain - எந்த Ross, Gordon மற்றும் Bagnell reduction ஆக frame செய்யுகிறார்கள் no-regret online learning க்கு.
அந்த framing விளைவு மக்கள் skip. Guarantee sequence of policies பற்றி iterations க்கு, single run அல்ல. இது பற்றி சொல்லுகிறது உங்கள் round three உதவினது. ஒரே வழி know செய்வது measure each iteration. Kelly மற்றும் colleagues introduced HG-DAgger ஏனெனில் classic DAgger கேட்கிறது expert க்கு action labels அப்போது novice still in control, எந்த paper வாதிட்டது என்பது decrease safety மற்றும், human experts உடன், degrade label quality மாயமான actuator lag மூலம் செய்யப்பட வாய்ப்பு உள்ளது. HG-DAgger also learns ஒரு safety threshold model-uncertainty-based risk metric க்கு மற்றும் reports improved performance இரண்டு DAgger மற்றும் behavioural cloning க்கு ஒரு driving task க்கு. இது publishes எதுவும் intervention counts அல்ல; அந்த உங்கள் produce yourself.
Rate வரையறுக்கும் இரண்டு மக்கள் அதே எண் get
Definition ஆக ஒரு line: intervention frames divided by run frames. Everything difficult hides என்ன counts ஆக frame மற்றும் என்ன counts ஆக run.
Frames, episodes அல்ல
Episodes counting at least one intervention useless ஆக: ten episodes ஒன்று nudge each மற்றும் ten இல் நீங்கள் drove eighty percent இரண்டும் give "10/10". Frame count separates அவர்கள், மற்றும் இது granularity recording ஏற்கனவே உள்ளது - LeRobot dataset format every timestep ஒரு row ஆக, எனவே intervention flag ஒரு boolean column state மற்றும் action க்கு அடுத்ததாக. இங்கே இது written per frame moment takeover starts மற்றும் cleared போது control returns.
Denominator corrections உள்ளடக்கிக்கொள்கிறது
இரண்டு denominators defensible ஆக - total frames of the run, அல்லது frames policy drove autonomously - மற்றும் அவர்கள் diverge badly high intervention levels. Take over 400 of 1000 frames மற்றும் first gives 40 percent, second 67 percent. Comparing one round measured first way against next measured second எவ்வாறு real improvement disappears. Take total frames மற்றும் never revisit choice mid-series.
Handover frames என்ன நிகழ்கிறது decide ஒரு சமயம்
Always seam உள்ளது. Control passes போது, சில frames belong neither side க்கு: arm held, leader aligning, recording frozen. Pipeline இல் அந்த handover frames stay raw stream மற்றும் never enter curated episode - அவர்கள் neither policy behaviour அல்லது correction training க்கு worth. Count seam same way every round: 30 Hz, six takeovers ஒரு two-second seam each 360 frames ஆக, enough move ஒரு percentage point.
Frames அல்லது episodes; total run அல்லது autonomous-only; handover frames in அல்லது out. Any of the three changed silently between rounds makes curve meaningless. Write all three down.
அதை logging எனவே number survives session
ஒரு metric status panel ஒரு running process readout ஆக: இது disappears restart மீது மற்றும் plotted முடியாது. One append-only line per run covers whole series - including which checkpoint நீங்கள் drove, அப்போது அது changes every round மற்றும் mixing அவர்கள் invalidates என்ன follows.
{"round": 2, "run_id": "inf-2026-08-24-1108", "checkpoint": "ckpt-8500",
"frames": 5412, "intervention_frames": 611, "rate": 0.113,
"takeovers": 7, "input": "keyboard", "denominator": "total_run_frames",
"handover_frames": "excluded", "episodes_kept_as_correction": 5,
"train_mix": {"base_demos": 120, "corrections": 41},
"eval_protocol": "protocol-A", "eval_trials": 20, "eval_successes": 11}இரண்டு fields carry weight. train_mix என்ன நீங்கள் fed next training job; without it nothing attributed முடியும். eval_protocol names fixed test நீங்கள் ran afterwards, மற்றும் field most often left empty ஆக. AY-Robots cockpit இல் takeover status reports intervention frame count running session, எனவே logging transcription rather than instrumentation; dataset documentation covers where per-frame columns land.

Curve படிக்கும்: மூன்று rounds, ஒரு verdict
மூன்று rounds minimum ஒரு reading, ஏனெனில் இரண்டு points always ஒரு line ஆக. Table below bookkeeping template placeholder numbers உடன், measurements from any run அல்ல. நீங்கள் தேடுகிறீர்கள் monotone decrease matched by increase success on fixed test.
| Round | Checkpoint driven | Frames | Intervention frames | Rate | Successes (of 20) |
|---|---|---|---|---|---|
| 0 (baseline) | base policy | 5 900 | 1 380 | 23.4 % | 7 |
| 1 | from round 0 mix | 5 610 | 980 | 17.5 % | 10 |
| 2 | from round 1 mix | 5 412 | 611 | 11.3 % | 11 |
| 3 | from round 2 mix | 5 380 | 598 | 11.1 % | 12 |
Rounds one மற்றும் two doing work ஆக. Round three அல்ல: 11.3 against 11.1 percent inside run-to-run variation anything measured physical arm மீது, மற்றும் fourth round same corrections spends afternoon nothing க்கு. Flat segment says change something structural.
- Remaining failures correctable அல்ல teleoperation மூலம் - object out of reach, gripper geometry, joint at its limit. Correction data fixes kinematic wall.
- Corrections too few base dataset against move gradient, மற்றும் nothing weights அவர்கள்.
- Corrections contradict each other, எனவே policy averages இரண்டு strategies மற்றும் lands between அவர்கள்.
- Failure upstream ஆக: camera moved, lighting changed, wrist view no longer matches training.
- Task at ceiling policy class இல் இந்த hardware மீது, மற்றும் next step more base data ஆக.
Last இரண்டு failures loop அல்ல. Sirius-Fleet இல் declining need human designed outcome ஆக: robot autonomy improves, anomaly predictors adapt prediction criteria, leading fewer requests human intervention மற்றும் gradually reducing human workload over time. There criterion adapts on purpose. Manual loop உங்கள் adapts too, silently - எனவே rate flattens ஏனெனில் task ceiling, looks exactly like one flattened ஏனெனில் operator stopped noticing. Which next problem ஆக.
Trap 1: falling rate rising success rate அல்ல
Intervention rate measures exactly one thing: எவ்வளவு run நீங்கள் decided own. அந்த decision உங்கள், made real time, மற்றும் இது moves. Round zero இல் நீங்கள் take over first bad approach angle; round three மூலம் நீங்கள் watched policy recover dozen அவர்களிலிருந்து மற்றும் நீங்கள் let it try. உங்கள் threshold moved; policy may not have. Rate falls either way.
Human trust at least modelled quantity literature rather than assumed constant ஆக. Sirius re-weights training samples approximated human trust உடன் மற்றும் optimises policy weighted behavioural cloning உடன், reporting 8 percent boost simulation மற்றும் 27 percent real hardware over state art policy success rate, at twice convergence speed. அது treats trust ஒரு weight data மீது. இது correct threshold உங்கள் head round zero மற்றும் round three க்கு இடையே அல்ல.
நீங்கள் remove drift முடியாது, எனவே pair rate something உங்கள் threshold touch முடியாது: fixed set trials இல் நீங்கள் intervene அல்ல, scored against criterion written before round. Rate falls அப்போது paired success rate stays put signature habituation - worth catching, ஏனெனில் loop இல் இருந்து inner இது feels like progress.
| Intervention rate | Success rate fixed protocol மீது | Most likely reading |
|---|---|---|
| falls | rises | Round worked. Continue. |
| flat | உங்கள் threshold drifted, அல்லது corrections removed effort failures remove இல்லாமல். | Corrections degrading policy. Check mixture மற்றும் அவர்களுடைய internal consistency. |
| Round bought nothing. Change mixture, checkpoint அல்லது task collecting more before. | Regression. Suspect training mix, changed camera அல்லது checkpoint mix-up before suspecting method. | Trap 2: fixed protocol இல்லாமல், நீங்கள் measure room |
| Real-robot evaluation expensive மற்றும் hard repeat. SIMPLER authors motivate simulated evaluation observation உடன் real-world evaluation such policies scalable அல்ல மற்றும் faces reproducibility challenges likely worsen policies broaden task spectrum. RoboArena attacks it other side, more than 600 pairwise real-robot evaluation episodes across seven generalist policies, run evaluators seven academic institutions DROID platform மீது. அவர்களுடைய evaluators pick அவர்களுடைய own tasks மற்றும் environments ஆனால் must judge pairs policies double-blind; paper reports அந்த crowd-sourced ranking tracks generalist-policy performance more accurately than conventional, centralised evaluation. | நீங்கள் run 600 paired episodes one மீது இல்லை | SO-100 |
| ஒரு workshop. என்ன நீங்கள் செய்ய முடியும் remove variance நீங்கள் control - list boring enough இது usually gets skipped. | Dimension | Freeze this |
என்ன drifts நீங்கள் செய்யவில்லை என்றால்
Start pose
ஒரு written home position, driven before every trialTrajectory starts out of distributionObjects
| Taped marks, மற்றும் same physical objects reserved evaluation க்கு | ஒரு 'better' policy really ஒரு closer object | Cameras மற்றும் light |
|---|---|---|
| Same mounts, indices, exposure; blinds closed; photograph setup | Swapped camera indices alone dominate result ஆக முடியும் | Trials மற்றும் stop rule |
| Fixed n, fixed timeout, criterion written before round | Post-hoc criteria turn near-misses into whatever நீங்கள் need | Operator behaviour |
| No interventions during evaluation trials | Evaluation becomes another correction session | பின்னர் arithmetic, unforgiving trial counts workshop afford முடியும். Normal approximation கீழ், 60 percent success over 20 trials carries standard error near 11 percentage points - square root 0.6 times 0.4 over 20 - மற்றும் difference between இரண்டு such rounds carries about 15. Jump 60 से 70 percent therefore consistent nothing having happened உடன். Fifty trials bring per-round figure roughly 7 points, மற்றும் cost afternoon. |
| இந்த asymmetry argument intervention rate க்கு: computed over thousands frames rather than twenty binary outcomes, இது moves earlier மற்றும் more smoothly. Caveat என்பது frames episode க்கு within heavily correlated - one bad grasp yields hundred consecutive intervention frames - எனவே effective sample size closer number takeovers. Treat rate early indicator மற்றும் success rate slow ground truth. | Keep evaluation episodes out training pool | Evaluation episodes must never enter composed training dataset - otherwise round n+1 scored data it trained, மற்றும் curve measures memorisation. |
| Trap 3: training only corrections மீது bends policy | Corrections interesting data, எனவே instinct train அவர்கள் மீது. செய்யாதீர்கள். அவர்கள் come by construction narrow slice state space policy ஏற்கனவே failing மற்றும் say almost nothing majority run worked. Fine-tune அந்த slice alone மற்றும் நீங்கள் get policy good recovering from botched approach forgotten எவ்வாறு clean ஒன்றை make. | Mandlekar மற்றும் colleagues built அவர்களுடைய remote intervention system around இந்த. அவர்களுடைய framing: manipulation tasks contain bottleneck regions requiring sequence precise actions - inserting pod into coffee machine is அவர்களுடைய example - where small deviations lead into states demonstrations never covered. அவர்களுடைய algorithm trains iteratively new data எனவே policy learns traverse அந்த bottlenecks, மற்றும் அவர்கள் report agents trained intervention data beat agents trained equivalent number samples from non-interventional demonstrators. |
Corrections-only fine-tuning versus composed mixture
என்ன corrections-only gets நீங்கள்
Fast: short run over few dozen episodes cheap enough repeat
Targeted: gradient dominated states நீங்கள் care
Cheap testing whether correction style learnable at all
Fine-tune distribution no longer resembles task distribution
- training documentation
- . என்ன tool உங்கள் செய்யவில்லை recording which ratio produced which curve.
- Trap 4: human consistent expert அல்ல
- DAgger theory assumes expert. நீங்கள் technical sense இல் ஒரு அல்ல: உங்கள் corrections samples fixed conditional distribution over actions அல்ல. ACT authors name இது when listing obstacles fine manipulation - errors compound over time, மற்றும் human demonstrations non-stationary ஆக முடியும். அவர்களுடைய system still reaches 80 to 90 percent success six real tasks from ten minutes demonstrations, எனவே problem tractable, absent அல்ல.
- Robomimic study makes point data side. Across six offline algorithms on five simulated மற்றும் three real-world multi-stage tasks, its lessons include sensitivity design choices, dependence demonstration quality, மற்றும் - one hurts most here - variability arising from stopping criteria, ஏனெனில் training மற்றும் evaluation objectives differ. Checkpoint lowest loss உடன் reliably highest success rate உடன் ஒன்று அல்ல.
- Hardware version இந்த உள்ளது: input mode shapes correction. ஒரு
covers என்ன each mode sends.
Weighting the interventions: IWR மற்றும் என்ன came after
Corrections valuable minority என்றால், principled fix weight rather than exclusivity. Intervention Weighted Regression reference implementation: data partitioned intervention மற்றும் non-intervention samples, மற்றும் இரண்டு partitions sampled equal proportion during training. Correction set five percent frames பின்னர் contributes half gradient, nominal behaviour vanishing இல்லாமல் distribution.
Equal proportion starting point, law அல்ல. Sirius generalises binary partition replacing continuous weight from approximated human trust; Sirius-Fleet moves decision upstream, using visual world model மற்றும் anomaly predictors decide when human asked at all. Common thread: intervention flag training signal, just bookkeeping column அல்ல.MethodHuman acts போது who decidesIntervention data எவ்வாறு usedReported result
DAgger (2011)
Fixed schedule; expert labels visited states
One growing dataset, unweighted
| Framed reduction no-regret online learning | HG-DAgger (2019) | Human, gating control real system மீது | Aggregated; plus learned safety threshold model uncertainty மீது |
|---|---|---|---|
| Better than DAgger மற்றும் BC driving மீது; intervention counts given எதுவும் இல்லை | IWR (2020) | Human, via remote teleoperation | Intervention மற்றும் non-intervention samples drawn equal proportion |
| Beats non-interventional demos equal sample count | Sirius (2022) | Human, deployment போது | Weighted BC, weights from approximated human trust |
| 8 % gain simulation, 27 % real hardware மீது; 2x faster convergence | ThriftyDAgger (2021) | System, gating novelty மற்றும் risk மீது | Interactive collection supervision budget கீழ் |
| User study (N=10): 58 % higher human மற்றும் 80 % higher robot performance than next best method | Fleet-DAgger (2022) | System, allocating attention fleet க்கு across | Fleet learning scored by Return Human Effort மீது |
| Up to 8.8x higher ROHE than baselines | Sirius-Fleet (2024) | Anomaly predictors self-adapting criteria உடன் | Multi-task learning visual world model உடன் |
| Fewer intervention requests autonomy improves போது | Leaderboard view comparing robot policies by measured score | Ranking policies against each other means something only when every entry ran same protocol. That applies உங்கள் own மூன்று rounds too. | நான்கு metrics worth logging next rate |
| Takeovers per run. Twenty short corrections மற்றும் one long one give similar rates மற்றும் describe different policies - one jittery, one single blind spot உடன். | Mean intervention length. Rising length falling count உடன் means remaining failures hard ones, which late progress looks like. | Time first intervention. Policy gets further before needing help improving even when total rate flat. | Where interventions cluster. Bin flag by normalised episode progress; stable peak across rounds points one bottleneck மீது. |

Freeze evaluation protocol before round zero
- Write down start pose, object placement, cameras, trial count, timeout மற்றும் success criterion. Photograph table. Document change to வேண்டும் என்றால், series restarts.
- Baseline அளவிடுக
- Run protocol interventions இல்லாமல் மற்றும் record successes; பின்னர் run ஒரு instrumented session takeover enabled உடன் மற்றும் record rate. அந்த இரண்டு numbers round zero ஆக.
- Collect corrections one consistent style உள்ளே
One input mode per round, one operator நீங்கள் manage முடிந்தால். Take over criterion நீங்கள் state out loud - 'gripper more than two centimetres off approach' - மற்றும் hold it round க்கு.
Triage every episode same day
- 1Correction, evaluation அல்லது discard. Ambiguous episodes get discarded, not saved theory on more data helps - correction நீங்கள் yourself fumbled ஒரு episode இல்லை worse.
Compose mixture explicitly
- 2Original demonstrations plus corrections, episode selection explicit per source. Record base count, correction count மற்றும் any weighting - இது field நீங்கள் want மூன்று weeks.
Continue checkpoint, பின்னர் re-measure
- 3Train composed dataset previous checkpoint from rather than base model, run frozen protocol plus one instrumented session, append row, மற்றும் compare against previous இரண்டு rounds.
இரண்டு limits change எவ்வாறு நீங்கள் read weak round. Continuing checkpoint initialises weights from it - not optimiser resume, எனவே schedule starts fresh மற்றும் short run converged checkpoint move very little முடியும். மற்றும் leader-align motion takeover start மீது has least mileage real hardware மீது anything chain, corrections all begin odd transient உடன் என்றால், look there before blaming mixture.
- 4Measured loop, already wired up
AY-Robots implements இந்த six steps product features ஆக: take over mid-run leader arm, keyboard அல்லது sliders, intervention frames flagged automatically உடன்; triage each episode correction, evaluation அல்லது discard ஆக; compose next dataset original demonstrations மற்றும் corrections, episode selection explicit per source இல் இருந்து; மற்றும் start next training run previous checkpoint instead base model.
- 5DAgger loop எவ்வாறு works பார்க்கவும்
என்ன numbers உங்களைக் கூற முடியாது
- 6None இந்த solved, மற்றும் tidy curve உங்களை persuade செய்யாது வேறு. Intervention rate measures joint system - policy, hardware, operator, room - மற்றும் attributes nothing its own. இது judge corrections were good முடியாது, மற்றும் இது fall happily task that got easier ஏனெனில் object drifted two centimetres closer three weeks.
என்ன இது செய்கிறது vague impression turn ஒரு column நீங்கள் argue முடியும். Alternative better metric அல்ல; இது மூன்று weeks collecting corrections curve உங்களுக்குக் கூற வேண்டும், round மூன்றுக்குப் பிறகு, collecting stop. Survey literature defines interactive imitation learning human feedback given intermittently during robot execution, allowing online improvement behaviour; family wide, மற்றும் arrangement works இல்லை number per round இல்லாமல். பார்க்கவும் also
the SO-100 imitation learning guide
the DAgger loop page
implemented pipeline க்கு.
Intervention rate just one minus success rate ஆக?இல்லை. Rate measures எவ்வளவு run நீங்கள் took; success rate measures task got done whether you இல்லாமல். Run succeed முடியும் 30 percent intervention rate உடன், மற்றும் run interventions இல்லாமல் fail outright முடியும். அவர்கள் differ noise, rate moves smoothly thousands correlated frames, success rate jumps between handful binary outcomes. Log both.Evaluation trials எவ்வளவு success rate mean anything க்கு தேவை?More than feels reasonable. Normal approximation கீழ், 20 trials true 60 percent success rate carry standard error near 11 percentage points, எனவே 10-point movement between rounds indistinguishable noise மூலம்; 50 trials bring figure roughly 7. நீங்கள் afford 50 முடியாது என்றால், keep trial count identical across rounds மற்றும் treat small movements inconclusive ஆக.என்ன mixing ratio demonstrations corrections க்கு நான் start வேண்டும்?Documented starting point IWR: partition data intervention மற்றும் non-intervention samples மற்றும் draw அவர்கள் equal proportion, எனவே corrections contribute half gradient however small fraction frames அவர்கள். உங்கள் setup weight sampling முடியாது என்றால், approximate it through episode counts composing dataset மற்றும் record ratio. Training corrections alone option rule out.My intervention rate went up after round. Round wasted ஆக?
Necessarily அல்ல, ஆனால் check boring explanations first: checkpoint நீங்கள் drove match ஒன்று நீங்கள் trained, camera index அல்லது mount change, object placement drift, correction style consistent ஆக. All நான்கு clean மற்றும் paired success rate also fell என்றால், suspect mixture - too few base demonstrations corrections against, அல்லது corrections contradict each other. Genuine regression belongs log, bin அல்ல.▾
Skip fixed evaluation protocol மற்றும் just watch intervention rate முடியும்?
Only நீங்கள் accept improvement habituation से tell முடியாது. Rate depends உங்கள் own real-time threshold stepping, மற்றும் அந்த threshold falls நீங்கள் used policy quirks. Frozen protocol part measurement உங்கள் threshold reach முடியாது - taped marks, written home pose, fixed trial count, criterion decided before round. இது stay unchanged series க்கு across.▾
Keep log repository, notebook இல் அல்ல: rounds days apart, hardware gets rebuilt, மற்றும் person reading curve October நீங்கள் August memory இல்லாமல்.
documentation FAQ▾
covers operational details left here.
My intervention rate went up after round. Round wasted ஆக?▾
Necessarily அல்ல, ஆனால் check boring explanations first: checkpoint நீங்கள் drove match ஒன்று நீங்கள் trained, camera index அல்லது mount change, object placement drift, correction style consistent ஆக. All நான்கு clean மற்றும் paired success rate also fell என்றால், suspect mixture - too few base demonstrations corrections against, அல்லது corrections contradict each other. Genuine regression belongs log, bin அல்ல.
Skip fixed evaluation protocol மற்றும் just watch intervention rate முடியும்?▾
Only நீங்கள் accept improvement habituation से tell முடியாது. Rate depends உங்கள் own real-time threshold stepping, மற்றும் அந்த threshold falls நீங்கள் used policy quirks. Frozen protocol part measurement உங்கள் threshold reach முடியாது - taped marks, written home pose, fixed trial count, criterion decided before round. இது stay unchanged series க்கு across.
Keep log repository, notebook இல் அல்ல: rounds days apart, hardware gets rebuilt, மற்றும் person reading curve October நீங்கள் August memory இல்லாமல். documentation FAQ covers operational details left here.
Sources
- Ross, Gordon & Bagnell (2011): A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger)
- Ross & Bagnell (2010): Efficient Reductions for Imitation Learning (AISTATS, PMLR v9)
- Kelly, Sidrane, Driggs-Campbell & Kochenderfer: HG-DAgger - Interactive Imitation Learning with Human Experts (arXiv 2018, ICRA 2019)
- Mandlekar et al. (2020): Human-in-the-Loop Imitation Learning using Remote Teleoperation
- IWR project page (Stanford): Intervention Weighted Regression
- Mandlekar et al. (2021): What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (robomimic)
- Liu, Nasiriany, Zhang, Bao & Zhu (2022): Robot Learning on the Job - Human-in-the-Loop Autonomy and Learning During Deployment (Sirius)
- Liu et al. (2024): Multi-Task Interactive Robot Fleet Learning with Visual World Models (Sirius-Fleet)
- Hoque et al. (2022): Fleet-DAgger - Interactive Robot Fleet Learning with Scalable Human Supervision
- Hoque et al. (2021): ThriftyDAgger - Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning
- Celemin et al. (2022): Interactive Imitation Learning in Robotics - A Survey
- Atreya et al. (2025): RoboArena - Distributed Real-World Evaluation of Generalist Robot Policies
- Li et al. (2024): Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER)
- Zhao, Kumar, Levine & Finn (2023): Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)
Ready for high-quality robotics data?
AY-Robots connects your robots to skilled operators worldwide.
Get Started