
व्यवहार क्लोनिंग विशेषज्ञ के अवस्था वितरण पर एक नीति को फिट करता है और फिर इसे स्वतंत्र रूप से तैनात किया जाता है। उन दोनों वितरणों के बीच का अंतर यह है कि क्यों एक नीति जो सत्यापन में ठीक दिखती है वह चरण 300 पर मेज से गिरती है। यह हमारी डीएजर श्रृंखला का सिद्धांत अध्याय है: द्विघातीय त्रुटि पद कहाँ से आता है, डेटासेट एकत्रीकरण क्या बदलता है, कोई-पश्चाताप प्रमाण क्या मानता है, और बिल का कौन सा हिस्सा मानव विशेषज्ञ को अभी भी भुगतान करना पड़ता है।
एक विशेष विफलता है जिससे हर कोई जो एक हेरफेर नीति को प्रशिक्षित करता है वह जल्दी या बाद में मिलता है। नीति घन के लिए पहुंचती है, दो सेंटीमीटर के भीतर आती है, हिचकिचाती है, बग़ल में विचलित होती है, फिर कार्य से असंबंधित कुछ करती है। सत्यापन loss ठीक था। आयोजित-बाहर एपिसोड के विरुद्ध open-loop प्रजनन ठीक था। और फिर भी आर्म एक ऐसी मुद्रा में समाप्त होता है जो प्रशिक्षण डेटा में कहीं नहीं दिखाई देती है, और वहाँ से इसके पास कहने के लिए कुछ समझदारी नहीं है।
उस विफलता का एक नाम है और इसके पीछे एक स्थिर सिद्धांत है। यह डीएजर पर चार लेख में से पहला है, और यह तर्क स्वयं को कवर करता है: क्यों प्रदर्शनकारी के अपने प्रक्षेपवक्र पर एक नीति को फिट करना एक त्रुटि पैदा करता है जो एपिसोड length के वर्ग के साथ बढ़ सकती है, डेटासेट एकत्रीकरण क्या बदलता है, और कोई-पश्चाताप प्रमाण क्या नहीं वचन देता है। real hardware पर डीएजर loop चलाने को SO-100 पर एक डीएजर loop चलाना में कवर किया गया है, मानव-gated variant को एचजी-डीएजर और मानव-gated हस्तक्षेप में, और माप प्रश्न को एक डीएजर loop को मापना में।
संक्षिप्त संस्करण
- •व्यवहार क्लोनिंग प्रदर्शनकारी के अवस्था वितरण पर प्रशिक्षित करता है और नीति के अपने पर मूल्यांकन किया जाता है। असमान एपिसोड के दौरान।
- •रॉस और बैगनेल दिखाते हैं कि अतिरिक्त लागत T वर्ग गुणा per-step त्रुटि के रूप में बढ़ सकती है; डीएजर paper इस सीमा को सुधारता है और नोट करता है कि यह तंग है।
- •डीएजर नीति स्वयं देखता है के states को लेबल करता है, और हर dataset एकत्रित किए गए पर retrain करता है, केवल नवीनतम को नहीं।
- •गारंटी एक reduction है no-regret online learning के लिए: एकत्रीकरण और retraining Follow-The-Leader है।
- •यह policy class में प्राप्य सर्वोत्तम loss के सापेक्ष है, शून्य के सापेक्ष नहीं - और विशेषज्ञ को अभी भी states को लेबल करना पड़ता है जो यह कभी नहीं बनाता।
व्यवहार क्लोनिंग चुपचाप मानती है
एक demonstration dataset एक pile है observation-action pairs की। Behavior cloning इस pile पर एक function को ordinary supervised learning के साथ फिट करता है और वहीं बंद हो जाता है। यह field में सबसे पुरानी विचार है। Pomerleau का ALVINN, 1988 में, एक three-layer back-propagation network था जो camera और laser range finder से images लेता था और direction पैदा करता था जो vehicle को travel करना चाहिए; यह simulated road images पर प्रशिक्षित किया गया था और कुछ field conditions में real roads का अनुसरण किया। Recipe ज्यादा नहीं बदली है; networks बदले हैं।
जो छोड़ा जाता है वह यह check है कि ये pairs कहाँ से आते हैं। उनमें से हर एक एक trajectory पर lie करता है जो demonstrator produced किया। Policy जो आप deploy करते हैं अपने खुद के produce करता है। जैसे ही यह deviate करता है, यह उन states के बारे में queried है जो training distribution में नहीं थे, और इसका answer इसे आगे out ले जाता है। Ross, Gordon और Bagnell डीएजर paper को बिल्कुल इससे खोलते हैं: sequential prediction i.i.d. assumption को violate करता है underlying statistical learning के, क्योंकि learner के स्वयं के predictions निर्धारित करते हैं कि inputs यह अगला कौन से देखता है।
उस paper में सबसे स्पष्ट illustration बिल्कुल robot नहीं है। Cloning एक near-optimal planner को Super Mario Bros के लिए एक policy produced की जो repeatedly एक obstacle के विरुद्ध फंस गई बजाय इसे jump करने के। Reason संपूर्ण argument है एक sentence में: expert हमेशा एक comfortable distance से jumped, तो dataset में कोई state नहीं था जिसमें Mario एक obstacle के विरुद्ध pressed up था, और इसलिए कोई label नहीं था क्या करना है एक बार वह था।
Mario को एक SO-100 arm के लिए swap करें और structure बिल्कुल समान है। आपके demonstrations एक clean approach और clean grasp दिखाते हैं, gripper closing से नहीं दो centimetres short - तो policy के पास कोई विचार नहीं है वहाँ से क्या करना है, और जो भी यह guesses इसे आगे out ले जाता है। Covariate shift एक property है डेटा संग्रह procedure का, network architecture का नहीं।
जहाँ द्विघातीय पद से आता है
2010 AISTATS paper Ross और Bagnell द्वारा, Efficient Reductions for Imitation Learning, compounding को precise बनाता है। T को task horizon बनें, task cost को unit interval में bounded बनें, और epsilon को surrogate loss बनें measured under expert's state distribution - number आपका validation set reports। फिर उस policy को T steps के लिए चलाने की extra cost, relative to expert, T squared times epsilon से bounded है। Ross, Gordon और Bagnell इसे Theorem 2.1 के रूप में डीएजर paper में restate करते हैं और sentence add करते हैं जो मायने रखता है: bound tight है। Problems exist करती हैं जहाँ एक policy जिसके पास epsilon loss है expert's distribution पर वास्तव में extra cost incur करता है जो T में quadratically बढ़ता है।
Tight मतलब typical नहीं है। Quadratic term एक worst case है over एक class of problems, नहीं एक prediction आपके pick-and-place task के लिए। What it establishes है कि more expert demonstration remove नहीं कर सकता है problem: यह केवल sharpen करता है estimate epsilon की एक distribution पर policy को tested नहीं किया जाएगा।
Escape route same paper में है, restated Theorem 2.2 के रूप में। अगर एक policy achieves loss epsilon under its own state distribution, और single wrong action costs most u में cost-to-go में under expert, extra cost bounded है u times T times epsilon द्वारा - linear in horizon। Constant u interesting quantity है: most 1 for 0-1 disagreement with expert, और O(1) जब भी expert कुछ steps में recover कर सकता है। Worst case में यह O(T) है, और linear bound फिर quadratic one से बेहतर नहीं है।
| Setting | Extra cost की bound expert पर | यह किस पर निर्भर करता है |
|---|---|---|
| Behavior cloning (Ross & Bagnell 2010, restated Thm. 2.1 में Ross et al. 2011) | T squared times epsilon | epsilon measured expert's state distribution पर; cost in [0,1]; bound tight है |
| Any policy जिसके पास epsilon loss है under its own distribution (Thm. 2.2) | u times T times epsilon | u bounds cost-to-go penalty का one wrong action; most 1 for 0-1 loss, O(T) worst case |
| Forward training (Ross & Bagnell 2010) | u times T times epsilon | one policy per timestep; T policies need करता है और known, finite T |
| SMILe (Ross & Bagnell 2010) | near-linear T और epsilon में some problem classes पर | alpha in O(1/T squared), N in O(T squared log T); yields stochastic mixture |
| डीएजर (Thm. 3.2, Ross et al. 2011) | u times T times epsilon_N, plus O(1) | N order of uT पर; strongly convex bounded loss; no-regret learner; epsilon_N सर्वश्रेष्ठ loss है hindsight में |

दो attempts जो आए थे डीएजर से पहले
Forward training honest लेकिन impractical answer है। Train separate policy हर timestep के लिए, in order, हर एक पर state distribution induced by policies पहले से ही fixed for earlier steps, तो हर policy देखता है exactly distribution यह face करेगा। Catch description में है: T policies, trained sequentially, no early stopping। Manipulation एपिसोड के लिए 30 frames per second पर, T hundreds में है।
SMILe, same paper से, और SEARN, Daume, Langford और Marcu के work से structured prediction पर, दूसरा route लेते हैं: एक stationary policy, लेकिन stochastic। हर iteration एक component train करता है और इसे mixture में add करता है, shifting probability mass away from expert। Result एक mixture है जिसमें कुछ components दूसरों से बदतर हैं - physical arm पर, एक controller जो bad component को mid-motion में sample कर सकता है। That stated motivation है wanting करने के लिए एक stationary deterministic policy इसके बजाय।
डीएजर: एक idea, एक box
Dataset Aggregation deterministic policy को keep करता है और fix को data collection में moves करता है। हर round: current policy को roll out करें, states record करें यह visits, expert पूछें क्या correct action होता आपके पास हर एक में, उन pairs को dataset add करें आपके पास पहले से, union पर retrain करें। Name algorithm है - आप aggregate करते हैं, आप कभी discard नहीं करते।
D <- {} # the aggregate dataset
pi_hat_1 <- any policy in Pi
for i = 1 .. N:
pi_i = beta_i * expert + (1 - beta_i) * pi_hat_i
roll out pi_i for T steps, record every visited state s
D_i = { (s, expert(s)) for every visited state s }
D = D union D_i # aggregate, do not replace
pi_hat_{i+1} = train on all of D
return the best pi_hat_i on a validation setतीन details बहुत अधिक weight carry करते हैं वह कैसे दिखते हैं से। Labels mixed policy द्वारा visited states के लिए हैं, लेकिन actions expert से आते हैं - policy questions provide करती है, expert answers। Retraining पूरे aggregate पर है, जो हर round को Follow-The-Leader step बनाता है: round n पर आप pick करते हैं सर्वश्रेष्ठ policy hindsight में over हर trajectory so far। That framing है क्या proof hangs पर। और algorithm sequence से सर्वश्रेष्ठ policy को return करके end करता है chosen on एक validation set, क्योंकि theorems guarantee करते हैं कि some policy sequence में अच्छी है, अंतिम एक नहीं।
बीटा अनुसूची, और क्यों यह एक tuning knob नहीं है
Mixed policy beta_i times expert प्लस एक minus beta_i times learner है। Point practical है: पहले कुछ learned policies बहुत कम data पर trained हैं, make कई mistakes, और अन्यथा rollout में spend करेंगे states जो irrelevant बन जाते हैं एक बार policy improves।
Theory exactly एक condition impose करता है: running average of betas को zero जाना चाहिए। Analysis काम करता है beta_i bounded के साथ (1 - alpha) को power i-1, एक constant alpha के लिए independent of T।
| Schedule | यह क्या करता है | Paper क्या reports करता है |
|---|---|---|
| beta_1 = 1 | First round pure expert demonstration है; no initial policy needed | Recommended starting point हर variant में |
| beta_i = 1 if i = 1, else 0 | Expert केवल round one में; कोई free parameter नहीं | Paper का parameter-free version, जो यह कहता है अक्सर सर्वश्रेष्ठ perform करता है practice में; 2980 Super Mario Bros पर 20 iterations के बाद |
| beta_i = p^(i-1) with p = 0.5 | Expert probability decay geometrically करती है | 3030 same benchmark पर, slightly ahead of parameter-free version |
| beta_i = p^(i-1) with p = 0.9 | Expert loop में far longer रहता है | Markedly slower convergence; still improving जब 20 iterations ended |
Gap 2980 और 3030 के बीच एक scale पर running to roughly 4300 छोटा है, लेकिन paper की explanation यह सबसे useful practical note है section में। Parameter-free schedule के साथ, Mario same spot में early stuck हो गया और mass of near-duplicate data generate किया उस एक location से; letting expert drive fraction of time दोनों unstuck किया उसे और widened variety of states। Schedule mixing ratio के बारे में कम है की तुलना में कि क्या आपके data collection नए states produce करना रखता है या same failure।
एक stochastic per-timestep mixture मतलब switching control authority at control rate, 30 times एक second पर typical SO-100 setup। No teleoperation interface बनाता है safe या meaningful। Real hardware पर beta schedule देता है way एक human decision को about when to take over: एक different algorithm एक different analysis के साथ।
Guarantee: एक reduction no-regret online learning को
यहाँ move है जो करता है paper को क्या है। Treat हर डीएजर round को एक example online learning problem में, जहाँ loss at round i surrogate loss है under state distribution of policy used at round i। Learner commits एक policy को before seeing उस loss, और sequence non-stationary है क्योंकि यह depend करता है policies पर produced so far।
एक algorithm no-regret है अगर its average loss over N rounds approaches कि सर्वश्रेष्ठ single policy का hindsight में। Follow-The-Leader on strongly convex losses ऐसी एक algorithm है, with average regret shrinking order of 1/N - और retraining पूरे aggregate पर precisely Follow-The-Leader है। Any other no-regret learner serve करेगा well: analysis एक reduction है, एक property नहीं एक optimiser का।
एक lemma gap को bridge करता है mixed policy के बीच जो collected data और learned policy जो will be deployed: Lemma 4.1 bounds L1 distance उनके state distributions के बीच by 2 T beta_i। यह क्यों है betas decay करने चाहिए - जब expert अभी भी appreciable control authority hold करता है, states आप collect करते हैं वह नहीं हैं states आपकी policy produce करेगी। Combine lemma को regret bound के साथ और main result follows: roughly T iterations के बाद, कुछ policy sequence में surrogate loss है under its own distribution within O(1/T) of epsilon_N। Feed उस को linear bound में और आप land करते हैं Theorem 3.2 पर।
Empirical side modest है by current standards। Super Tux Kart में supervised baseline अपनी average falls improve नहीं किया per lap as more data arrived, डीएजर reached एक policy जो कभी fell नहीं track से after fifteen iterations, और SMILe after twenty still fell roughly twice per lap। None of these एक manipulation result है।
क्या proof नहीं promise करता
Theorem statements conditional हैं, और conditions load-bearing हैं।
- एक bound linear राथर quadratic T में, stated assumptions के तहत।
- एक stationary deterministic policy राथर stochastic mixture।
- एक genuine reduction: any no-regret online learner slots में।
- एक concrete iteration count - roughly T rounds before regret term stops mattering।
- एक guarantee at least एक policy के लिए sequence में, hence closing validation pass।
- यह relative है epsilon_N को, सर्वश्रेष्ठ loss class में hindsight में, not to zero। अगर आपकी class नहीं represent कर सकती है expert, यह empty है practice में।
- यह needs एक no-regret method या एक strongly convex surrogate loss - stronger than classification reductions यह build पर, as authors note।
- Constant u O(T) हो सकता है worst case में, और linear bound फिर collapses back to quadratic।
- यह bounds iterations, not expert labels। Robot पर, labels हैं budget।
- यह assumes expert को query किया जा सकता है हर visited state पर और answers correctly वहाँ। यह assumption है entire cost।
एक further result अक्सर quoted है एक refutation के रूप में और नहीं है एक। Rajaraman, Yang, Jiao और Ramachandran study minimax limits of imitation learning in episodic MDPs with finite state space S और horizon H, और prove suboptimality lower bound on order of |S| H squared over N जो holds even जब learner may actively query expert at visited states। That एक worst-case rate है over एक class of MDPs at fixed episode budget, और क्या rules out है idea कि interaction improves minimax rate; डीएजर का theorem एक different statement है, bounding deployed policy को relative to क्या its own policy class achieve कर सकता है।
Swamy, Choudhury, Bagnell और Wu later classified ये algorithms by कौन से moments expert's behaviour के वह match करते हैं, और introduced notion of moment recoverability जो delineates कैसे well हर family compounding error को mitigates। Surveys by Osa और by Celemin cover algorithmic landscape और human-feedback interfaces।
Bill: labelling states expert कभी नहीं produced
हर thing उपरोक्त assumes एक expert जो query किया जा सकता है anywhere। Simulation में एक planner के साथ जो nearly free है - Mario experiments used एक near-optimal planner with full access to game state। With एक human on एक robot यह dominant cost है, और peculiar एक: human को produce करना पड़ता है एक correct action में एक configuration जिसका own competence कभी नहीं बनाया है।
Kelly, Sidrane, Driggs-Campbell और Kochenderfer state objection directly HG-डीएजर paper में। Vanilla डीएजर requires expert supply करने को action labels जबकि fully in control नहीं हो रहा system का। यह reduces safety, और with human experts यह likely है degrade करने को quality of collected labels, जो वह put down करते हैं to perceived actuator lag। Label आप get back करते हैं नहीं है label algorithm assumed।
Laskey और colleagues attack problem को other side से with DART, और उनकी framing blunt है: on-policy techniques tedious हैं human supervisors के लिए, add computational burden, और may visit dangerous states during training। उनके alternative injects calibrated noise into supervisor के own demonstrations, तो recovery gets demonstrated without robot कभी running एक untrusted policy। MuJoCo Humanoid पर वह report करते हैं DART decreasing supervisor का cumulative reward by 5 percent during training, जबकि डीएजर executes policies with 80 percent less cumulative reward than supervisor; grasping in clutter में एक Toyota HSR के साथ, average 62 percent increase over behavior cloning।
Zhang और Cho के SafeDAgger treats queries को reference policy को as scarce resource: एक separate safety policy predicts, without querying, क्या primary policy about है to deviate from reference beyond threshold, और only उन states handed over हैं। सभी तीन react करते हैं same fact को - डीएजर analysis charges कुछ नहीं expert labels के लिए, और reality charges बहुत।
Labelling off-distribution states mentally harder है than demonstrating task। एक normal demonstration मतलब executing एक motor plan आपके पास already है। Correcting एक policy जो put किया है gripper को somewhere आप कभी नहीं करते हैं मतलब constructing एक recovery on spot, under time pressure, with robot still moving। Expect fewer usable minutes per session than plain recording session में, और watch अपने own correction quality decay करते हैं over course एक का।

क्या यह मतलब है एक SO-100 के लिए आपकी desk पर
Translate horizon को अपने own units में। एक twenty-second episode at 30 frames per second है 600 decision steps, और T हर bound में above है कि number। At T = 600, difference एक term scaling के बीच T और एक scaling with T squared है difference एक policy के बीच जो recover करता है एक bad approach से और एक जो नहीं करता।
यह part है क्यों action chunking helps: जब एक policy emits एक short sequence actions per inference step का, number decision points drops, और तो does number chances to compound। Zhao, Kumar, Levine और Finn name compounding error as motivation for Action Chunking with Transformers, और report 80 से 90 percent success on six difficult real-world tasks, on low-cost bimanual hardware, from ten minutes worth of demonstrations। Chunking remove नहीं करता covariate shift - states still हैं policy's own - लेकिन shortens effective horizon। देखें action chunking और SO-100 अनुकरण सीखने guide।
दूसरा translation है progress metric। आप measure नहीं कर सकते epsilon under policy's own distribution directly - जो needs है ground-truth expert actions हर visited state के लिए, thing आप हैं trying करने करना avoid producing। क्या एक human-gated loop आपको देता है instead है intervention rate: fraction frames में एक run के दौरान जिसमें human had taken over। यह proxy है, और यह moves for reasons unrelated to policy - एक patient operator intervenes कम। Used consistently, यह one number है जो says क्या एक round था worth afternoon।
एक तीसरा translation है एक data-quality warning analysis cover नहीं करता। Mandlekar और colleagues studied छह offline learning algorithms on पाँच simulated and तीन real-world multi-stage manipulation tasks, और report sensitivity to algorithmic design choices, dependence on quality of demonstrations, और variability caused by stopping criterion। Belkhale, Cui और Sadigh argue कि dataset quality should be formalized through action divergence और transition diversity, और note कि state diversity हमेशा beneficial नहीं है। एक डीएजर round adds states nobody chose deliberately: कुछ हैं recovery data आपको need, कुछ हैं robot flailing जबकि आप fumble करते हैं takeover control के लिए।
Mechanically एक round है छह steps: run inference with recording on, take over जब policy misbehaves, review run और file हर episode, sync corrections, compose mixed dataset से originals plus corrections with episode selection made explicitly per source, और continue training से previous checkpoint राथर base model से। ay-robots पर those steps exist as buttons, जो removes plumbing लेकिन नहीं judgement। दो caveats: continuing from checkpoint initializes weights और नहीं है एक optimizer resume, और leader-arm alignment move still है lightly tested on hardware। देखें training और datasets।
डीएजर loop, पहले से ही wired up
Takeover during live inference run, per-frame intervention marking, filing episodes as corrections या evaluations, composing mixed dataset with explicit episode selection per source, और continuing training from existing checkpoint सभी built in हैं। आप अभी भी decide करते हैं जब take over करना है और क्या keep करना है - वह part automate नहीं करता।
देखें डीएजर loop कैसे काम करता हैFamily tree, एक table में
| Method | कौन choose करता है states | क्या expert supplies करता है | Main cost |
|---|---|---|---|
| Behavior cloning | Expert | Clean demonstrations | कोई recovery data नहीं; error T में quadratically compound कर सकता है |
| Forward training | Learner, per timestep | Labels induced distribution के साथ | T separate policies; long horizons के लिए unusable |
| SMILe / SEARN | Expert और learner का stochastic mixture | Labels mixture's distribution के साथ | Mixture के components quality में differ करते हैं |
| डीएजर | Mixed policy, बीटा decay करता है zero को | Correct action हर visited state के लिए | Labelling states expert कभी produce नहीं करते, while नहीं control में |
| DART | Expert, perturbed by injected noise | Demonstrations under calibrated noise | Noise must be calibrated learner के error को |
| एचजी-डीएजर | Learner, जब तक human take नहीं करता over | Corrections only in human-gated segments | Depends on human का judgement about जब intervene करना है |
| SafeDAgger | Learner, filtered by safety gate | Labels केवल जब gate ask करता है | Gate itself must be trained और trusted |
अक्सर पूछे जाने वाले प्रश्न
क्या मैं actually observe करूँगा quadratic error growth अपने robot पर?▾
Clean curve के रूप में नहीं। Bound worst case है: tight में कि कुछ problem attain करता है, not कि आपका करेगा। क्या देखते हैं held-out frames पर एक consequence है - एक policy जो scores well, real task पर fails, और improve नहीं करता जब अधिक record करते हैं same का। अगर अधिक clean data helping बंद करता है, वह है covariate shift, एक data-volume problem नहीं।
क्या मुझे implement करना पड़ता है बीटा mixture को call करने के लिए यह डीएजर?▾
Parameter-free version - expert round one में, pure learner afterwards - एक legitimate special case है often performed best में original experiments। क्या drop नहीं कर सकते है aggregation: retraining केवल newest corrections पर breaks Follow-The-Leader interpretation, जो है जहाँ no-regret argument से आता है। Training corrections केवल एक पर है बहुत weaker procedure।
क्यों return करें सर्वश्रेष्ठ policy validation set पर instead of last एक?▾
क्योंकि theorems guarantee एक good policy exists कहीं sequence में, नहीं कि यह है final iterate - bound है minimum over sequence पर। Shipping जो भी came out of last round discards एक stated condition result का, और last round नहीं है reliably सर्वश्रेष्ठ।
कितने rounds plan करने चाहिए मैं?▾
Theory wants iterations order of T, जो 600-step episode के लिए है एक number anyone नहीं run करता hardware पर। Original experiments ran twenty iterations हर benchmark पर। Practice में आप run करते rounds जब तक intervention rate बंद नहीं होती falling, far नीचे count analysis assumes - एक real gap between theory और practice।
क्या अगर मेरी policy class simply नहीं कर सकती represent expert?▾
फिर डीएजर आपको save नहीं करता, और bound कहता है - यह expressed है relative to epsilon_N, सर्वश्रेष्ठ loss class में hindsight। अगर वह है large क्योंकि wrong architecture का, missing observation या camera जो देख नहीं सकता scene, aggregation आपको देता है एक policy जो optimal है within एक class जो नहीं कर सकता task। Run open-loop replay held-out episodes के विरुद्ध before आप collect corrections।
यहाँ कहाँ जाएँ से
अगर आपने train नहीं किया एक policy अभी, यह theory है premature: record एक dataset पहले, शुरुआत से अपनी first policy training करना और desktop client। अगर आप weighing हैं एक और सौ clean demonstrations के विरुद्ध शुरुआत करना corrections: clean demonstrations नहीं करते fix distribution problem। Mechanics के लिए, continue करें human-gated variant और फिर SO-100 walkthrough।
Sources
- Ross & Bagnell (2010): Efficient Reductions for Imitation Learning (AISTATS, PMLR v9)
- Ross, Gordon & Bagnell (2011): A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
- Ross, Gordon & Bagnell (2011), AISTATS proceedings version (PMLR v15, pp. 627-635)
- Pomerleau (1988): ALVINN - An Autonomous Land Vehicle in a Neural Network (NeurIPS)
- Daume III, Langford & Marcu (2009): Search-based Structured Prediction (SEARN)
- Laskey, Lee, Fox, Dragan & Goldberg (2017): DART - Noise Injection for Robust Imitation Learning
- Kelly, Sidrane, Driggs-Campbell & Kochenderfer (2018): HG-DAgger - Interactive Imitation Learning with Human Experts
- Zhang & Cho (2016): Query-Efficient Imitation Learning for End-to-End Autonomous Driving (SafeDAgger)
- Osa, Pajarinen, Neumann, Bagnell, Abbeel & Peters (2018): An Algorithmic Perspective on Imitation Learning
- Celemin et al. (2022): Interactive Imitation Learning in Robotics - A Survey
- Rajaraman, Yang, Jiao & Ramachandran (2020): Toward the Fundamental Limits of Imitation Learning
- Swamy, Choudhury, Bagnell & Wu (2021): Of Moments and Matching - A Game-Theoretic Framework for Closing the Imitation Gap
- Mandlekar et al. (2021): What Matters in Learning from Offline Human Demonstrations for Robot Manipulation (robomimic)
- Zhao, Kumar, Levine & Finn (2023): Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)
- Belkhale, Cui & Sadigh (2023): Data Quality in Imitation Learning (NeurIPS)
Ready for high-quality robotics data?
AY-Robots connects your robots to skilled operators worldwide.
Get Started