Tabletop robot workspace நிலையான evaluation setup ஆக பயன்படுத்தப்படும் வகை repeated policy runs க்கு
DAggerமதிப்பீடுImitation LearningRobot தரவுSO-100

DAgger Loop ஐ அளவிடுதல்: Intervention Rate, Evaluation Protocol, மற்றும் இடையே உள்ள பொறிகள்

AY-Robots ResearchAugust 27, 202615 நிமிட வாசனை

Intervention rate - intervention frames ஐ run frames மூலம் வகுப்பதால் - மனிதனால் கட்டுப்படுத்தப்படும் DAgger loops மிக விலையுயர்ந்த நேர்மையான முன்னேற்ற சமிக்ஞை. இந்த கட்டுரை அதை வரையறுக்கிறது, இரண்டு மனிதர்களும் அதே மதிப்பை கணக்கிட பயன்படுத்தப்படுகிறது, அதை எவ்வாறு பதிவு செய்வது என்பதைக் காட்டுகிறது, மற்றும் நான்கு வழிகளை வேலை செய்கிறது அது உங்களை வழிநடத்துவது: operator habituation, ஒரு drifting evaluation setup, corrections மீது மட்டும் பயிற்சி, மற்றும் ஒரு மனிதன் செவ்வாய்கிழமையிலும் திங்கட்கிழமையிலும் வேறுபட்ட முறையில் திருத்துகிறான்.

DAgger round உணர்வு உபயுக்த நீங்கள் அதில் இருக்கும் போது. நீங்கள் policy ஐ இயக்குகிறீர்கள், grasp ஐ தவறவிட்டபோது எடுத்துக் கொள்ளுங்கள், corrections ஐ சேமிக்கவும், retrain செய்யவும், அடுத்த run - நீங்கள் சொல்வதுண்டு - கொஞ்சம் மெருமெரு என்று. இரண்டு rounds பின்னர் எதுவும் மாற்றியுள்ளதா என்று நீங்கள் கூற முடியாது, ஏனெனில் "கொஞ்சம் மெருமெரு" ஒரு அளவு அல்ல.

Loop ஒரு round க்கு ஒரு எண்ணுக்கு தேவை, மற்றும் ஒரு human-gated loop இல் அந்த எண் கிட்டத்தட்ட இலவசம்: run இன் பகுதி நீங்கள் policy க்கு பதிலாக control இருந்தீர்கள். இந்த கட்டுரை அதை வரையறுக்கிறது, இரண்டு மனிதர்களும் அதே மதிப்பை கணக்கிட பயன்படுத்தப்படுகிறது, அதை பதிவுசெய்யுங்கள், விளைந்த curve ஐ படிக்கவும், மற்றும் நான்கு வழிகளை பெயரிடுங்கள் இது உங்களை வழிநடத்துகிறது எப்போது இது தனிப்பட்டதாக நிற்கிறது. இந்த பகுதி கருதும் கோட்பாடுக்கு, பார்க்கவும் the DAgger explainer மற்றும் the human-gated variant; hands-on counterpart ஆக running a DAgger loop on an SO-100.

Short version

  • Intervention rate = intervention frames / run frames. Frames, episodes அல்ல, மற்றும் denominator நீங்கள் correcting செலவழிச்ச frames உள்ளடக்கியது.
  • மூன்று rounds க்கு ஒரு flat curve அர்த்தம் rounds எதுவும் வாங்கவில்லை - mixture, checkpoint அல்லது task ஐ நான்காவது collect செய்வதற்கு பதிலாக மாற்றவும்.
  • Intervention rate falling ஆக இருப்பது rising success rate அல்ல. Stepping in க்கான உங்கள் threshold drift downward என்று நம்பிக்கை வளரும் போது.
  • Frozen evaluation protocol இல்லாமல் - same start poses, objects, cameras, trial count - நீங்கள் room ஐ அளவிடுகிறீர்கள்.
  • Corrections மீது மட்டும் training state distribution ஐ skew செய்கிறது. IWR இன் பதில்: intervention மற்றும் non-intervention samples சமம் விகிதம் draw செய்யவும்.

ஏன் round க்கு ஒரு எண் தேவை

DAgger உள்ளது distribution problem காரணத்தால், மற்றும் distribution problems கண்ணுக்கு கணக்கிடப்பட முடியாது. Ross மற்றும் Bagnell காட்டினர் imitation learning fixed demonstration set மீது leading compounding errors மற்றும் regret bound growing quadratically time horizon இல்: demonstrator visited states கள் கூட learner கூட என்பது ஒரு சமயம், data இல் எதுவும் மீண்டும் வர எவ்வாறு பற்றி சொல்லுகிறது. DAgger fixes இதை iterating மூலம் - run the policy, label the states it visits, aggregate, retrain - எந்த Ross, Gordon மற்றும் Bagnell reduction ஆக frame செய்யுகிறார்கள் no-regret online learning க்கு.

அந்த framing விளைவு மக்கள் skip. Guarantee sequence of policies பற்றி iterations க்கு, single run அல்ல. இது பற்றி சொல்லுகிறது உங்கள் round three உதவினது. ஒரே வழி know செய்வது measure each iteration. Kelly மற்றும் colleagues introduced HG-DAgger ஏனெனில் classic DAgger கேட்கிறது expert க்கு action labels அப்போது novice still in control, எந்த paper வாதிட்டது என்பது decrease safety மற்றும், human experts உடன், degrade label quality மாயமான actuator lag மூலம் செய்யப்பட வாய்ப்பு உள்ளது. HG-DAgger also learns ஒரு safety threshold model-uncertainty-based risk metric க்கு மற்றும் reports improved performance இரண்டு DAgger மற்றும் behavioural cloning க்கு ஒரு driving task க்கு. இது publishes எதுவும் intervention counts அல்ல; அந்த உங்கள் produce yourself.

Rate வரையறுக்கும் இரண்டு மக்கள் அதே எண் get

Definition ஆக ஒரு line: intervention frames divided by run frames. Everything difficult hides என்ன counts ஆக frame மற்றும் என்ன counts ஆக run.

Frames, episodes அல்ல

Episodes counting at least one intervention useless ஆக: ten episodes ஒன்று nudge each மற்றும் ten இல் நீங்கள் drove eighty percent இரண்டும் give "10/10". Frame count separates அவர்கள், மற்றும் இது granularity recording ஏற்கனவே உள்ளது - LeRobot dataset format every timestep ஒரு row ஆக, எனவே intervention flag ஒரு boolean column state மற்றும் action க்கு அடுத்ததாக. இங்கே இது written per frame moment takeover starts மற்றும் cleared போது control returns.

Denominator corrections உள்ளடக்கிக்கொள்கிறது

இரண்டு denominators defensible ஆக - total frames of the run, அல்லது frames policy drove autonomously - மற்றும் அவர்கள் diverge badly high intervention levels. Take over 400 of 1000 frames மற்றும் first gives 40 percent, second 67 percent. Comparing one round measured first way against next measured second எவ்வாறு real improvement disappears. Take total frames மற்றும் never revisit choice mid-series.

Handover frames என்ன நிகழ்கிறது decide ஒரு சமயம்

Always seam உள்ளது. Control passes போது, சில frames belong neither side க்கு: arm held, leader aligning, recording frozen. Pipeline இல் அந்த handover frames stay raw stream மற்றும் never enter curated episode - அவர்கள் neither policy behaviour அல்லது correction training க்கு worth. Count seam same way every round: 30 Hz, six takeovers ஒரு two-second seam each 360 frames ஆக, enough move ஒரு percentage point.

மூன்று decisions break comparability

Frames அல்லது episodes; total run அல்லது autonomous-only; handover frames in அல்லது out. Any of the three changed silently between rounds makes curve meaningless. Write all three down.

அதை logging எனவே number survives session

ஒரு metric status panel ஒரு running process readout ஆக: இது disappears restart மீது மற்றும் plotted முடியாது. One append-only line per run covers whole series - including which checkpoint நீங்கள் drove, அப்போது அது changes every round மற்றும் mixing அவர்கள் invalidates என்ன follows.

json
{"round": 2, "run_id": "inf-2026-08-24-1108", "checkpoint": "ckpt-8500",
 "frames": 5412, "intervention_frames": 611, "rate": 0.113,
 "takeovers": 7, "input": "keyboard", "denominator": "total_run_frames",
 "handover_frames": "excluded", "episodes_kept_as_correction": 5,
 "train_mix": {"base_demos": 120, "corrections": 41},
 "eval_protocol": "protocol-A", "eval_trials": 20, "eval_successes": 11}
One line per run. Recording denominator மற்றும் handover convention அடுத்ததாக number matters more than field names.

இரண்டு fields carry weight. train_mix என்ன நீங்கள் fed next training job; without it nothing attributed முடியும். eval_protocol names fixed test நீங்கள் ran afterwards, மற்றும் field most often left empty ஆக. AY-Robots cockpit இல் takeover status reports intervention frame count running session, எனவே logging transcription rather than instrumentation; dataset documentation covers where per-frame columns land.

Training configuration matrix showing policies, datasets and checkpoints side by side
Every round changes checkpoint மற்றும் dataset mix. ஒரு number whose combination recorded இல்லை attributed முடியாது.

Curve படிக்கும்: மூன்று rounds, ஒரு verdict

மூன்று rounds minimum ஒரு reading, ஏனெனில் இரண்டு points always ஒரு line ஆக. Table below bookkeeping template placeholder numbers உடன், measurements from any run அல்ல. நீங்கள் தேடுகிறீர்கள் monotone decrease matched by increase success on fixed test.

RoundCheckpoint drivenFramesIntervention framesRateSuccesses (of 20)
0 (baseline)base policy5 9001 38023.4 %7
1from round 0 mix5 61098017.5 %10
2from round 1 mix5 41261111.3 %11
3from round 2 mix5 38059811.1 %12

Rounds one மற்றும் two doing work ஆக. Round three அல்ல: 11.3 against 11.1 percent inside run-to-run variation anything measured physical arm மீது, மற்றும் fourth round same corrections spends afternoon nothing க்கு. Flat segment says change something structural.

  • Remaining failures correctable அல்ல teleoperation மூலம் - object out of reach, gripper geometry, joint at its limit. Correction data fixes kinematic wall.
  • Corrections too few base dataset against move gradient, மற்றும் nothing weights அவர்கள்.
  • Corrections contradict each other, எனவே policy averages இரண்டு strategies மற்றும் lands between அவர்கள்.
  • Failure upstream ஆக: camera moved, lighting changed, wrist view no longer matches training.
  • Task at ceiling policy class இல் இந்த hardware மீது, மற்றும் next step more base data ஆக.

Last இரண்டு failures loop அல்ல. Sirius-Fleet இல் declining need human designed outcome ஆக: robot autonomy improves, anomaly predictors adapt prediction criteria, leading fewer requests human intervention மற்றும் gradually reducing human workload over time. There criterion adapts on purpose. Manual loop உங்கள் adapts too, silently - எனவே rate flattens ஏனெனில் task ceiling, looks exactly like one flattened ஏனெனில் operator stopped noticing. Which next problem ஆக.

Trap 1: falling rate rising success rate அல்ல

Intervention rate measures exactly one thing: எவ்வளவு run நீங்கள் decided own. அந்த decision உங்கள், made real time, மற்றும் இது moves. Round zero இல் நீங்கள் take over first bad approach angle; round three மூலம் நீங்கள் watched policy recover dozen அவர்களிலிருந்து மற்றும் நீங்கள் let it try. உங்கள் threshold moved; policy may not have. Rate falls either way.

Human trust at least modelled quantity literature rather than assumed constant ஆக. Sirius re-weights training samples approximated human trust உடன் மற்றும் optimises policy weighted behavioural cloning உடன், reporting 8 percent boost simulation மற்றும் 27 percent real hardware over state art policy success rate, at twice convergence speed. அது treats trust ஒரு weight data மீது. இது correct threshold உங்கள் head round zero மற்றும் round three க்கு இடையே அல்ல.

நீங்கள் remove drift முடியாது, எனவே pair rate something உங்கள் threshold touch முடியாது: fixed set trials இல் நீங்கள் intervene அல்ல, scored against criterion written before round. Rate falls அப்போது paired success rate stays put signature habituation - worth catching, ஏனெனில் loop இல் இருந்து inner இது feels like progress.

Intervention rateSuccess rate fixed protocol மீதுMost likely reading
fallsrisesRound worked. Continue.
flatஉங்கள் threshold drifted, அல்லது corrections removed effort failures remove இல்லாமல்.Corrections degrading policy. Check mixture மற்றும் அவர்களுடைய internal consistency.
Round bought nothing. Change mixture, checkpoint அல்லது task collecting more before.Regression. Suspect training mix, changed camera அல்லது checkpoint mix-up before suspecting method.Trap 2: fixed protocol இல்லாமல், நீங்கள் measure room
Real-robot evaluation expensive மற்றும் hard repeat. SIMPLER authors motivate simulated evaluation observation உடன் real-world evaluation such policies scalable அல்ல மற்றும் faces reproducibility challenges likely worsen policies broaden task spectrum. RoboArena attacks it other side, more than 600 pairwise real-robot evaluation episodes across seven generalist policies, run evaluators seven academic institutions DROID platform மீது. அவர்களுடைய evaluators pick அவர்களுடைய own tasks மற்றும் environments ஆனால் must judge pairs policies double-blind; paper reports அந்த crowd-sourced ranking tracks generalist-policy performance more accurately than conventional, centralised evaluation.நீங்கள் run 600 paired episodes one மீது இல்லை SO-100
ஒரு workshop. என்ன நீங்கள் செய்ய முடியும் remove variance நீங்கள் control - list boring enough இது usually gets skipped.DimensionFreeze this

என்ன drifts நீங்கள் செய்யவில்லை என்றால்

Start pose

ஒரு written home position, driven before every trialTrajectory starts out of distributionObjects

Taped marks, மற்றும் same physical objects reserved evaluation க்குஒரு 'better' policy really ஒரு closer objectCameras மற்றும் light
Same mounts, indices, exposure; blinds closed; photograph setupSwapped camera indices alone dominate result ஆக முடியும்Trials மற்றும் stop rule
Fixed n, fixed timeout, criterion written before roundPost-hoc criteria turn near-misses into whatever நீங்கள் needOperator behaviour
No interventions during evaluation trialsEvaluation becomes another correction sessionபின்னர் arithmetic, unforgiving trial counts workshop afford முடியும். Normal approximation கீழ், 60 percent success over 20 trials carries standard error near 11 percentage points - square root 0.6 times 0.4 over 20 - மற்றும் difference between இரண்டு such rounds carries about 15. Jump 60 से 70 percent therefore consistent nothing having happened உடன். Fifty trials bring per-round figure roughly 7 points, மற்றும் cost afternoon.
இந்த asymmetry argument intervention rate க்கு: computed over thousands frames rather than twenty binary outcomes, இது moves earlier மற்றும் more smoothly. Caveat என்பது frames episode க்கு within heavily correlated - one bad grasp yields hundred consecutive intervention frames - எனவே effective sample size closer number takeovers. Treat rate early indicator மற்றும் success rate slow ground truth.Keep evaluation episodes out training poolEvaluation episodes must never enter composed training dataset - otherwise round n+1 scored data it trained, மற்றும் curve measures memorisation.
Trap 3: training only corrections மீது bends policyCorrections interesting data, எனவே instinct train அவர்கள் மீது. செய்யாதீர்கள். அவர்கள் come by construction narrow slice state space policy ஏற்கனவே failing மற்றும் say almost nothing majority run worked. Fine-tune அந்த slice alone மற்றும் நீங்கள் get policy good recovering from botched approach forgotten எவ்வாறு clean ஒன்றை make.Mandlekar மற்றும் colleagues built அவர்களுடைய remote intervention system around இந்த. அவர்களுடைய framing: manipulation tasks contain bottleneck regions requiring sequence precise actions - inserting pod into coffee machine is அவர்களுடைய example - where small deviations lead into states demonstrations never covered. அவர்களுடைய algorithm trains iteratively new data எனவே policy learns traverse அந்த bottlenecks, மற்றும் அவர்கள் report agents trained intervention data beat agents trained equivalent number samples from non-interventional demonstrators.

Corrections-only fine-tuning versus composed mixture

என்ன corrections-only gets நீங்கள்

என்ன இது costs

Fast: short run over few dozen episodes cheap enough repeat

Targeted: gradient dominated states நீங்கள் care

Cheap testing whether correction style learnable at all

Fine-tune distribution no longer resembles task distribution

Nominal behaviour degrade visibly can within single round
Rate fall while success falls அதனுடன் - clean segments traded recoveries
  • training documentation
  • . என்ன tool உங்கள் செய்யவில்லை recording which ratio produced which curve.
  • Trap 4: human consistent expert அல்ல
Mixing ratio hyperparameter மற்றும் deserves written down. Composing next dataset where it decided: original demonstrations plus correction set, episode selection explicit per source rather than rule quietly pulls in whatever available. Here அது ஒரு compose step cockpit, மற்றும் result ordinary dataset - பார்க்கவும்
  • DAgger theory assumes expert. நீங்கள் technical sense இல் ஒரு அல்ல: உங்கள் corrections samples fixed conditional distribution over actions அல்ல. ACT authors name இது when listing obstacles fine manipulation - errors compound over time, மற்றும் human demonstrations non-stationary ஆக முடியும். அவர்களுடைய system still reaches 80 to 90 percent success six real tasks from ten minutes demonstrations, எனவே problem tractable, absent அல்ல.
  • Robomimic study makes point data side. Across six offline algorithms on five simulated மற்றும் three real-world multi-stage tasks, its lessons include sensitivity design choices, dependence demonstration quality, மற்றும் - one hurts most here - variability arising from stopping criteria, ஏனெனில் training மற்றும் evaluation objectives differ. Checkpoint lowest loss உடன் reliably highest success rate உடன் ஒன்று அல்ல.
  • Hardware version இந்த உள்ளது: input mode shapes correction. ஒரு

leader-follower setup produces continuous, human-paced trajectories. Keyboard nudges clamped hard server மூலம் - two degrees per call arm joints மீது, four gripper மீது - எனவே same intent arrives ஒரு staircase small steps. Sliders send absolute targets மற்றும் server moves at most six degrees toward அவர்கள் per call. மூன்று modes, மூன்று action distributions; mixing all மூன்று into one correction set மற்றும் பின்னர் wondering why approach turned twitchy self-inflicted ஆக. teleoperation documentation

covers என்ன each mode sends.

Weighting the interventions: IWR மற்றும் என்ன came after

Corrections valuable minority என்றால், principled fix weight rather than exclusivity. Intervention Weighted Regression reference implementation: data partitioned intervention மற்றும் non-intervention samples, மற்றும் இரண்டு partitions sampled equal proportion during training. Correction set five percent frames பின்னர் contributes half gradient, nominal behaviour vanishing இல்லாமல் distribution.

Equal proportion starting point, law அல்ல. Sirius generalises binary partition replacing continuous weight from approximated human trust; Sirius-Fleet moves decision upstream, using visual world model மற்றும் anomaly predictors decide when human asked at all. Common thread: intervention flag training signal, just bookkeeping column அல்ல.MethodHuman acts போது who decidesIntervention data எவ்வாறு usedReported result

DAgger (2011)

Fixed schedule; expert labels visited states

One growing dataset, unweighted

Framed reduction no-regret online learningHG-DAgger (2019)Human, gating control real system மீதுAggregated; plus learned safety threshold model uncertainty மீது
Better than DAgger மற்றும் BC driving மீது; intervention counts given எதுவும் இல்லைIWR (2020)Human, via remote teleoperationIntervention மற்றும் non-intervention samples drawn equal proportion
Beats non-interventional demos equal sample countSirius (2022)Human, deployment போதுWeighted BC, weights from approximated human trust
8 % gain simulation, 27 % real hardware மீது; 2x faster convergenceThriftyDAgger (2021)System, gating novelty மற்றும் risk மீதுInteractive collection supervision budget கீழ்
User study (N=10): 58 % higher human மற்றும் 80 % higher robot performance than next best methodFleet-DAgger (2022)System, allocating attention fleet க்கு acrossFleet learning scored by Return Human Effort மீது
Up to 8.8x higher ROHE than baselinesSirius-Fleet (2024)Anomaly predictors self-adapting criteria உடன்Multi-task learning visual world model உடன்
Fewer intervention requests autonomy improves போதுLeaderboard view comparing robot policies by measured scoreRanking policies against each other means something only when every entry ran same protocol. That applies உங்கள் own மூன்று rounds too.நான்கு metrics worth logging next rate
Takeovers per run. Twenty short corrections மற்றும் one long one give similar rates மற்றும் describe different policies - one jittery, one single blind spot உடன்.Mean intervention length. Rising length falling count உடன் means remaining failures hard ones, which late progress looks like.Time first intervention. Policy gets further before needing help improving even when total rate flat.Where interventions cluster. Bin flag by normalised episode progress; stable peak across rounds points one bottleneck மீது.
ஒரு round protocol நீங்கள் actually run முடியும்
ஒரு measured round has six steps மற்றும் produces one row log இல். First நான்கு loop ஆக; last இரண்டு make it measurement.

Freeze evaluation protocol before round zero

  • Write down start pose, object placement, cameras, trial count, timeout மற்றும் success criterion. Photograph table. Document change to வேண்டும் என்றால், series restarts.
  • Baseline அளவிடுக
  • Run protocol interventions இல்லாமல் மற்றும் record successes; பின்னர் run ஒரு instrumented session takeover enabled உடன் மற்றும் record rate. அந்த இரண்டு numbers round zero ஆக.
  • Collect corrections one consistent style உள்ளே

One input mode per round, one operator நீங்கள் manage முடிந்தால். Take over criterion நீங்கள் state out loud - 'gripper more than two centimetres off approach' - மற்றும் hold it round க்கு.

Triage every episode same day

  1. 1
    Correction, evaluation அல்லது discard. Ambiguous episodes get discarded, not saved theory on more data helps - correction நீங்கள் yourself fumbled ஒரு episode இல்லை worse.

    Compose mixture explicitly

  2. 2
    Original demonstrations plus corrections, episode selection explicit per source. Record base count, correction count மற்றும் any weighting - இது field நீங்கள் want மூன்று weeks.

    Continue checkpoint, பின்னர் re-measure

  3. 3
    Train composed dataset previous checkpoint from rather than base model, run frozen protocol plus one instrumented session, append row, மற்றும் compare against previous இரண்டு rounds.

    இரண்டு limits change எவ்வாறு நீங்கள் read weak round. Continuing checkpoint initialises weights from it - not optimiser resume, எனவே schedule starts fresh மற்றும் short run converged checkpoint move very little முடியும். மற்றும் leader-align motion takeover start மீது has least mileage real hardware மீது anything chain, corrections all begin odd transient உடன் என்றால், look there before blaming mixture.

  4. 4
    Measured loop, already wired up

    AY-Robots implements இந்த six steps product features ஆக: take over mid-run leader arm, keyboard அல்லது sliders, intervention frames flagged automatically உடன்; triage each episode correction, evaluation அல்லது discard ஆக; compose next dataset original demonstrations மற்றும் corrections, episode selection explicit per source இல் இருந்து; மற்றும் start next training run previous checkpoint instead base model.

  5. 5
    DAgger loop எவ்வாறு works பார்க்கவும்

    என்ன numbers உங்களைக் கூற முடியாது

  6. 6
    None இந்த solved, மற்றும் tidy curve உங்களை persuade செய்யாது வேறு. Intervention rate measures joint system - policy, hardware, operator, room - மற்றும் attributes nothing its own. இது judge corrections were good முடியாது, மற்றும் இது fall happily task that got easier ஏனெனில் object drifted two centimetres closer three weeks.

    என்ன இது செய்கிறது vague impression turn ஒரு column நீங்கள் argue முடியும். Alternative better metric அல்ல; இது மூன்று weeks collecting corrections curve உங்களுக்குக் கூற வேண்டும், round மூன்றுக்குப் பிறகு, collecting stop. Survey literature defines interactive imitation learning human feedback given intermittently during robot execution, allowing online improvement behaviour; family wide, மற்றும் arrangement works இல்லை number per round இல்லாமல். பார்க்கவும் also

the SO-100 imitation learning guide

base dataset உங்கள் mix against க்கு,

running a policy

inference side க்கு, மற்றும்

the DAgger loop page

implemented pipeline க்கு.

Intervention rate just one minus success rate ஆக?இல்லை. Rate measures எவ்வளவு run நீங்கள் took; success rate measures task got done whether you இல்லாமல். Run succeed முடியும் 30 percent intervention rate உடன், மற்றும் run interventions இல்லாமல் fail outright முடியும். அவர்கள் differ noise, rate moves smoothly thousands correlated frames, success rate jumps between handful binary outcomes. Log both.Evaluation trials எவ்வளவு success rate mean anything க்கு தேவை?More than feels reasonable. Normal approximation கீழ், 20 trials true 60 percent success rate carry standard error near 11 percentage points, எனவே 10-point movement between rounds indistinguishable noise மூலம்; 50 trials bring figure roughly 7. நீங்கள் afford 50 முடியாது என்றால், keep trial count identical across rounds மற்றும் treat small movements inconclusive ஆக.என்ன mixing ratio demonstrations corrections க்கு நான் start வேண்டும்?Documented starting point IWR: partition data intervention மற்றும் non-intervention samples மற்றும் draw அவர்கள் equal proportion, எனவே corrections contribute half gradient however small fraction frames அவர்கள். உங்கள் setup weight sampling முடியாது என்றால், approximate it through episode counts composing dataset மற்றும் record ratio. Training corrections alone option rule out.My intervention rate went up after round. Round wasted ஆக?

Necessarily அல்ல, ஆனால் check boring explanations first: checkpoint நீங்கள் drove match ஒன்று நீங்கள் trained, camera index அல்லது mount change, object placement drift, correction style consistent ஆக. All நான்கு clean மற்றும் paired success rate also fell என்றால், suspect mixture - too few base demonstrations corrections against, அல்லது corrections contradict each other. Genuine regression belongs log, bin அல்ல.

Skip fixed evaluation protocol மற்றும் just watch intervention rate முடியும்?

Only நீங்கள் accept improvement habituation से tell முடியாது. Rate depends உங்கள் own real-time threshold stepping, மற்றும் அந்த threshold falls நீங்கள் used policy quirks. Frozen protocol part measurement உங்கள் threshold reach முடியாது - taped marks, written home pose, fixed trial count, criterion decided before round. இது stay unchanged series க்கு across.

Keep log repository, notebook இல் அல்ல: rounds days apart, hardware gets rebuilt, மற்றும் person reading curve October நீங்கள் August memory இல்லாமல்.

documentation FAQ

covers operational details left here.

My intervention rate went up after round. Round wasted ஆக?

Necessarily அல்ல, ஆனால் check boring explanations first: checkpoint நீங்கள் drove match ஒன்று நீங்கள் trained, camera index அல்லது mount change, object placement drift, correction style consistent ஆக. All நான்கு clean மற்றும் paired success rate also fell என்றால், suspect mixture - too few base demonstrations corrections against, அல்லது corrections contradict each other. Genuine regression belongs log, bin அல்ல.

Skip fixed evaluation protocol மற்றும் just watch intervention rate முடியும்?

Only நீங்கள் accept improvement habituation से tell முடியாது. Rate depends உங்கள் own real-time threshold stepping, மற்றும் அந்த threshold falls நீங்கள் used policy quirks. Frozen protocol part measurement உங்கள் threshold reach முடியாது - taped marks, written home pose, fixed trial count, criterion decided before round. இது stay unchanged series க்கு across.

Keep log repository, notebook இல் அல்ல: rounds days apart, hardware gets rebuilt, மற்றும் person reading curve October நீங்கள் August memory இல்லாமல். documentation FAQ covers operational details left here.

Ready for high-quality robotics data?

AY-Robots connects your robots to skilled operators worldwide.

Get Started