Montgomery et al., 2017 - Google Patents
Reset-free guided policy search: Efficient deep reinforcement learning with stochastic initial statesMontgomery et al., 2017
View PDF- Document ID
- 561241442632502531
- Author
- Montgomery W
- Ajay A
- Finn C
- Abbeel P
- Levine S
- Publication year
- Publication venue
- 2017 IEEE International Conference on Robotics and Automation (ICRA)
External Links
Snippet
Autonomous learning of robotic skills can allow general-purpose robots to learn wide behavioral repertoires without extensive manual engineering. However, robotic skill learning must typically make trade-offs to enable practical real-world learning, such as requiring …
- 230000002787 reinforcement 0 title abstract description 19
Classifications
-
- G—PHYSICS
- G06—COMPUTING; CALCULATING; COUNTING
- G06N—COMPUTER SYSTEMS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computer systems based on biological models
- G06N3/02—Computer systems based on biological models using neural network models
- G06N3/04—Architectures, e.g. interconnection topology
-
- G—PHYSICS
- G06—COMPUTING; CALCULATING; COUNTING
- G06N—COMPUTER SYSTEMS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N99/00—Subject matter not provided for in other groups of this subclass
- G06N99/005—Learning machines, i.e. computer in which a programme is changed according to experience gained by the machine itself during a complete run
-
- G—PHYSICS
- G06—COMPUTING; CALCULATING; COUNTING
- G06N—COMPUTER SYSTEMS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computer systems based on biological models
- G06N3/02—Computer systems based on biological models using neural network models
- G06N3/08—Learning methods
-
- G—PHYSICS
- G05—CONTROLLING; REGULATING
- G05B—CONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
- G05B13/00—Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion
- G05B13/02—Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric
- G05B13/0265—Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric the criterion being a learning criterion
- G05B13/027—Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric the criterion being a learning criterion using neural networks only
-
- G—PHYSICS
- G06—COMPUTING; CALCULATING; COUNTING
- G06K—RECOGNITION OF DATA; PRESENTATION OF DATA; RECORD CARRIERS; HANDLING RECORD CARRIERS
- G06K9/00—Methods or arrangements for reading or recognising printed or written characters or for recognising patterns, e.g. fingerprints
- G06K9/62—Methods or arrangements for recognition using electronic means
- G06K9/6217—Design or setup of recognition systems and techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
- G06K9/6232—Extracting features by transforming the feature space, e.g. multidimensional scaling; Mappings, e.g. subspace methods
- G06K9/6247—Extracting features by transforming the feature space, e.g. multidimensional scaling; Mappings, e.g. subspace methods based on an approximation criterion, e.g. principal component analysis
-
- G—PHYSICS
- G05—CONTROLLING; REGULATING
- G05B—CONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
- G05B17/00—Systems involving the use of models or simulators of said systems
- G05B17/02—Systems involving the use of models or simulators of said systems electric
-
- G—PHYSICS
- G05—CONTROLLING; REGULATING
- G05B—CONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
- G05B13/00—Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion
- G05B13/02—Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric
- G05B13/04—Adaptive control systems, i.e. systems automatically adjusting themselves to have a performance which is optimum according to some preassigned criterion electric involving the use of models or simulators
-
- G—PHYSICS
- G06—COMPUTING; CALCULATING; COUNTING
- G06N—COMPUTER SYSTEMS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computer systems utilising knowledge based models
-
- G—PHYSICS
- G06—COMPUTING; CALCULATING; COUNTING
- G06K—RECOGNITION OF DATA; PRESENTATION OF DATA; RECORD CARRIERS; HANDLING RECORD CARRIERS
- G06K9/00—Methods or arrangements for reading or recognising printed or written characters or for recognising patterns, e.g. fingerprints
- G06K9/62—Methods or arrangements for recognition using electronic means
- G06K9/6288—Fusion techniques, i.e. combining data from various sources, e.g. sensor fusion
- G06K9/629—Fusion techniques, i.e. combining data from various sources, e.g. sensor fusion of extracted features
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Montgomery et al. | Reset-free guided policy search: Efficient deep reinforcement learning with stochastic initial states | |
| Gu et al. | Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates | |
| Yahya et al. | Collective robot reinforcement learning with distributed asynchronous guided policy search | |
| Kahn et al. | Plato: Policy learning using adaptive trajectory optimization | |
| Chebotar et al. | Closing the sim-to-real loop: Adapting simulation randomization with real world experience | |
| Vecerik et al. | A practical approach to insertion with variable socket position using deep reinforcement learning | |
| Kaushik et al. | Fast online adaptation in robotics through meta-learning embeddings of simulated priors | |
| Jain et al. | Learning deep visuomotor policies for dexterous hand manipulation | |
| CN113677485B (en) | Efficient Adaptation of Robot Control Strategies for New Tasks Using Meta-learning Based on Meta-imitation Learning and Meta-reinforcement Learning | |
| Laskey et al. | Comparing human-centric and robot-centric sampling for robot deep learning from demonstrations | |
| Laskey et al. | Dart: Noise injection for robust imitation learning | |
| Zhang et al. | Learning deep neural network policies with continuous memory states | |
| US11403513B2 (en) | Learning motor primitives and training a machine learning system using a linear-feedback-stabilized policy | |
| Finn et al. | Deep visual foresight for planning robot motion | |
| Fu et al. | One-shot learning of manipulation skills with online dynamics adaptation and neural network priors | |
| Sæmundsson et al. | Meta reinforcement learning with latent variable gaussian processes | |
| Amarjyoti | Deep reinforcement learning for robotic manipulation-the state of the art | |
| Ren et al. | Adaptsim: Task-driven simulation adaptation for sim-to-real transfer | |
| Si et al. | Agen: Adaptable generative prediction networks for autonomous driving | |
| Bischoff et al. | Policy search for learning robot control using sparse data | |
| Fanger et al. | Gaussian processes for dynamic movement primitives with application in knowledge-based cooperation | |
| Tschiatschek et al. | Variational inference for data-efficient model learning in pomdps | |
| CN119501923A (en) | Human-in-the-loop tasks and motion planning in imitation learning | |
| Alt et al. | Robot program parameter inference via differentiable shadow program inversion | |
| Torabi et al. | Sample-efficient adversarial imitation learning from observation |