Deprecated: The each() function is deprecated. This message will be suppressed on further calls in /home/zhenxiangba/zhenxiangba.com/public_html/phproxy-improved-master/index.php on line 456
CN121223793B - Self-adaptive strategy optimization method and system for body robot based on interactive feedback - Google Patents
[go: Go Back, main page]

CN121223793B - Self-adaptive strategy optimization method and system for body robot based on interactive feedback - Google Patents

Self-adaptive strategy optimization method and system for body robot based on interactive feedback

Info

Publication number
CN121223793B
CN121223793B CN202511533027.1A CN202511533027A CN121223793B CN 121223793 B CN121223793 B CN 121223793B CN 202511533027 A CN202511533027 A CN 202511533027A CN 121223793 B CN121223793 B CN 121223793B
Authority
CN
China
Prior art keywords
feedback
strategy
robot
standardized
generate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Active
Application number
CN202511533027.1A
Other languages
Chinese (zh)
Other versions
CN121223793A (en
Inventor
张云飞
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Yunzhou Hongye Beijing Technology Co ltd
Original Assignee
Yunzhou Hongye Beijing Technology Co ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Yunzhou Hongye Beijing Technology Co ltd filed Critical Yunzhou Hongye Beijing Technology Co ltd
Priority to CN202511533027.1A priority Critical patent/CN121223793B/en
Publication of CN121223793A publication Critical patent/CN121223793A/en
Application granted granted Critical
Publication of CN121223793B publication Critical patent/CN121223793B/en
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Landscapes

  • Feedback Control In General (AREA)

Abstract

The invention discloses a self-adaptive strategy optimization method and a self-adaptive strategy optimization system for a robot with a body based on interactive feedback, and relates to the technical field of robots. The method comprises the steps of collecting operation data of the robot with the body in a task execution process, analyzing environment information, user voice, gestures and binary evaluation signals to generate standardized feedback characterization, calculating a reward function based on the environment information and the feedback characterization and generating a standardized reward value, forming a state-feedback joint representation through high-order mapping, generating a candidate strategy set, screening update strategy parameters under the constraint of strategy consistency and execution stability, generating an action sequence for the update parameters to execute operation, and forming consistency indexes and closed-loop data streams through comparison of feedback results and the standardized feedback characterization, so that strategy self-adaptive optimization is achieved. The intelligent generation, execution consistency and task stability improvement of the physical robot strategy are realized through the high-order fusion and closed-loop self-adaptive optimization of multi-mode perception and user interaction feedback.

Description

Self-adaptive strategy optimization method and system for body robot based on interactive feedback
Technical Field
The invention relates to the technical field of robots, in particular to a self-adaptive strategy optimization method and system for a self-adaptive strategy of a robot body based on interactive feedback.
Background
In practical application, the robot needs to interact with a complex dynamic environment and a user, and the accuracy and the intelligence of task execution are affected by various factors. In the prior art, an autonomous robot generally relies on a single sensory data or a predefined action strategy to accomplish the operation, which strategy generates multiple dependent static rules or low-order feedback control methods. This approach has various problems in complex environments:
on one hand, the conventional strategy cannot fully consider multi-mode feedback generated when the robot performs actions in the environment, so that response to the environment state and user intention is lagged or inaccurate in the strategy execution process. Particularly when force sense, spatial attitude and complex operation tasks are involved, it is difficult for a robot to realize high-precision motion control and continuous execution.
On the other hand, the existing method lacks a systematic closed-loop mechanism in the strategy optimization process, depends on single feedback or simple rewarding evaluation, and cannot form effective self-adaptive adjustment. The non-linear relation of the user interaction intention and the environment constraint is lack of modeling and quantification, so that the problems of strategy deviation, unstable execution, inconsistent behaviors and the like of the robot in the multi-task and multi-environment operation are easy to occur.
In addition, the conventional method has limited capabilities for fusion of multi-modal data and modeling of higher-order dependency relationships. The image, depth, force sense and pose data are often processed by low-order linearity, high-order nonlinear interaction between the environmental state information and user feedback is ignored, and the strategy optimization process lacks refinement, individuation and self-adaption capability.
Therefore, a technical scheme capable of fully utilizing multi-mode sensing data, analyzing user interaction feedback in real time and realizing strategy self-adaptive optimization through closed-loop data flow is needed, so that task execution capacity, feedback consistency and strategy intellectualization level of the robot in a dynamic complex environment are improved.
Disclosure of Invention
Based on the above-mentioned shortcomings of the prior art, the present invention aims to provide an adaptive strategy optimization method and system for an autonomous robot based on interactive feedback, so as to solve the above-mentioned technical problems.
In order to achieve the above purpose, the invention provides the following technical scheme that the self-adaptive strategy optimization method of the robot with body based on interactive feedback comprises the following steps:
Acquiring image data, depth data, force sense data and pose data of the robot body in the task execution process, and generating environment state information, wherein the environment state information comprises the current spatial pose, perception characteristics and stress state of the robot body;
Analyzing and structuring the environment state information, a voice command, a gesture command and a binary evaluation signal from a user to generate a standardized feedback representation, and expressing the joint characteristics of the environment state and the user interaction intention;
calculating a reward function according to the environmental state information and the standardized feedback characterization, generating a standardized reward value, and quantifying the consistency relation between the environmental state information and the standardized feedback characterization;
Performing high-order mapping on the environmental state information and the standardized reward value to generate a state-feedback joint representation, and reflecting the interaction relation of the state constraint and the feedback constraint in the strategy generation process;
Generating a candidate strategy set based on the state-feedback joint representation, and screening and generating updated strategy parameters from the candidate strategy set under constraint conditions, wherein the constraint conditions comprise strategy consistency constraint and execution stability constraint;
Mapping the updated strategy parameters into a robot action sequence, and performing operation and environment interaction by the robot according to the action sequence to generate a feedback execution result;
Comparing a feedback execution result generated after the robot executes the action sequence with a standardized feedback characterization to generate a consistency index, judging whether strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, generating a closed-loop data stream by the feedback execution result and the consistency index, and supporting self-adaptive optimization of strategies.
The invention is further configured that the collecting data of the robot with body in the task execution process, and generating the environmental status information includes:
In the task execution process, acquiring image data, depth data, force sense data and pose data of the robot body, preprocessing the acquired data, and generating a unified data structure;
performing multi-scale feature extraction on the image data to generate visual perception feature representation;
gradient coding and neighborhood feature mapping are carried out on the depth data, and space geometric feature representation is generated;
Performing joint stress weighting and nonlinear mapping on the force sense data to generate force sense state representation;
position coding and direction coding are carried out on pose data, and a spatial pose representation is generated;
And uniformly mapping and fusing the visual perception feature representation, the space geometric feature representation, the force sense state representation and the space posture representation to generate environment state information, wherein the environment state information comprises the current space posture, the perception feature and the stress state of the robot.
The invention is further arranged such that the generating a normalized feedback representation comprises:
Analyzing and structuring the environmental state information, the voice command, the gesture command and the binary evaluation signal of the user;
performing time sequence feature extraction and nonlinear coding on the voice command to generate voice feature representation;
mapping space and channel characteristics of the gesture instruction to generate gesture characteristic representation;
performing nonlinear embedding and weighting processing on the binary evaluation signals to generate evaluation characteristic representations;
and performing high-order nonlinear joint mapping on the voice feature representation, the gesture feature representation and the evaluation feature representation and the environment state information to generate a standardized feedback representation, and expressing joint features of the environment state and the user interaction intention.
The invention is further arranged such that the generating a normalized prize value comprises:
nonlinear mapping and high-order interaction are carried out on the environmental state information and the standardized feedback characterization, so that a state-feedback coupling characteristic is formed;
Nonlinear high-order integration is carried out on the state-feedback coupling characteristic, and a reward value is generated;
and carrying out normalization processing on the generated reward value to generate a standardized reward value.
The invention is further arranged for generating a state-feedback joint representation comprising:
Performing nonlinear transformation on the environmental state information, and strengthening state constraint characteristics;
nonlinear mapping is carried out on the standardized reward value, and feedback constraint characteristics are enhanced;
and carrying out high-order interaction on the mapped environmental state characteristics and the mapped standardized reward values to form state-feedback combined characteristics.
The invention further provides that the filtering and generating updated strategy parameters from the candidate strategy set under the constraint condition comprises the following steps:
Generating a candidate policy set based on the state-feedback joint representation;
performing policy consistency constraint evaluation on each candidate policy, and quantifying the consistency of the candidate policies in the continuous execution process;
performing stability constraint evaluation on each candidate strategy, and quantifying the robustness of the candidate strategy in a dynamic environment;
And screening strategies meeting the strategy consistency constraint and the execution stability constraint from the candidate strategy set, and generating updated strategy parameters.
The invention further provides that the generating feedback execution result comprises:
converting the updated strategy parameters into a robot action sequence, and executing operation and environment interaction by the robot according to the action sequence;
In the executing process of the robot action, the environment state changes along with the higher-order influence of the action sequence, and the environment state after the action is executed is generated;
And generating a feedback execution result according to the environment state and the action sequence after the action is executed, and quantifying the relation between the action and the environment response.
The method further provides that the determining whether the policy parameter accords with the feedback constraint according to the consistency index and the preset threshold value comprises the following steps:
Performing high-order difference calculation on a feedback execution result generated after the robot executes the action sequence and a standardized feedback characterization, and quantifying nonlinear deviation between the action sequence and the user interaction intention;
Generating a consistency index according to the high-order difference calculation result, and quantifying the consistency degree of strategy parameters and feedback constraints;
Judging whether the strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, and forming a strategy parameter closed-loop judgment result;
And combining the feedback execution result with the consistency index to generate a closed-loop data stream.
The invention is further arranged that the generating of the closed-loop data stream, the adaptive optimization of the support strategy comprises:
according to the closed-loop data stream, high-order error extraction is carried out on a feedback execution result, a consistency index and a strategy parameter judgment result after the robot executes an action sequence;
Generating nonlinear optimization increment according to the high-order error, and performing self-adaptive adjustment on strategy parameters to form an updated strategy parameter set;
And generating a next round of robot action sequence according to the updated strategy parameters, and realizing the consistency of the action sequence with the user interaction intention and the environment constraint.
The invention also provides an adaptive strategy optimization system of the self-adaptive robot based on the interactive feedback, which comprises:
The multi-mode sensing data acquisition module acquires image data, depth data, force sense data and pose data of the robot body in the task execution process, and generates environment state information, wherein the environment state information comprises the current spatial pose, sensing characteristics and stress state of the robot body;
The interactive feedback analysis module analyzes and constructs the environment state information, the voice command, the gesture command and the binary evaluation signal from the user to generate a standardized feedback representation, and expresses the joint characteristics of the environment state and the interactive intention of the user;
the instant rewarding modeling module calculates a rewarding function according to the environment state information and the standardized feedback characterization, generates a standardized rewarding value and quantifies the consistency relation between the environment state information and the standardized feedback characterization;
The state-feedback joint mapping module is used for carrying out high-order mapping on the environmental state information and the standardized rewarding value to generate a state-feedback joint representation, and reflecting the interaction relation of the state constraint and the feedback constraint in the strategy generation process;
The constraint optimization strategy generation module generates a candidate strategy set based on the state-feedback joint representation, screens and generates updated strategy parameters from the candidate strategy set under constraint conditions, wherein the constraint conditions comprise strategy consistency constraint and execution stability constraint;
the action reasoning and control mapping module maps the updated strategy parameters into a robot action sequence, and the robot performs operation and environment interaction according to the action sequence to generate a feedback execution result;
And the execution feedback consistency verification module compares an execution result after the robot executes the action sequence with the standardized feedback characterization to generate a consistency index, judges whether strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, generates a closed-loop data stream according to the execution result and the consistency index, and supports self-adaptive optimization of the strategy.
The invention provides a self-adaptive strategy optimization method and system of a self-adaptive robot based on interactive feedback, wherein the method comprises the steps of acquiring image data, depth data, force sense data and pose data of the self-adaptive robot in a task execution process to generate environment state information, wherein the environment state information comprises the current spatial pose, perception characteristics and stress state of the self-adaptive robot, analyzing and structuring the environment state information and voice instructions, gesture instructions and binary evaluation signals from a user to generate standardized feedback characterization, expressing joint characteristics of the environment state and the user interactive intention, calculating a reward function according to the environment state information and the standardized feedback characterization, generating a standardized reward value, quantifying the consistency relation between the environment state information and the standardized feedback characterization, performing high-order mapping on the environment state information and the standardized reward value, generating a state-feedback joint representation, reflecting the interactive relation between the state constraint and the feedback constraint in a strategy generation process, generating a candidate strategy set based on the state-feedback joint representation, screening and generating updated strategy parameters under constraint conditions from the candidate strategy set, wherein the constraint conditions comprise consistency constraint and execution stability constraint, generating a sequence of the updated strategy parameters and the execution stability constraint parameters, generating a self-adaptive strategy performance index according to a preset strategy, comparing the updated strategy sequence with a preset strategy execution sequence with a preset motion index, and a self-adaptive performance index of the self-adaptive strategy optimization system of the self-adaptive robot, generating a self-adaptive strategy optimization system, and a self-adaptive strategy optimization method based on the sequence of the performance index is obtained by comparing the sequence of the performance of the robot with the performance of the performance with a preset sequence, the beneficial effects include:
1. the strategy self-adaptability and execution consistency are improved, namely, the consistency quantification of environment state information and user intention is realized through joint modeling of multi-mode perception data and user interaction feedback, and the strategy parameters are subjected to high-order self-adaptive optimization by combining closed-loop data flow, so that the robot with the body can continuously execute tasks under a dynamic complex environment and maintain the high consistency of strategy execution;
2. The multi-mode information fusion and high-order interaction capability is enhanced, namely, high-order fusion and nonlinear mapping are carried out on image, depth, force sense and pose data, and simultaneously, voice, gesture and binary evaluation signals are structured to form standardized feedback characterization, so that high-order interaction of state constraint and feedback constraint in the generation process of a robot strategy is realized, and the strategy generation is more refined and intelligent;
3. and the task execution stability and the robustness are improved, namely policy consistency constraint and execution stability constraint are introduced when a candidate policy set is generated, the policy parameters are adjusted through high-order nonlinear optimization increment, and the action execution stability and the robustness to environmental disturbance of the robot under a complex environment are obviously improved by combining closed-loop consistency verification.
The foregoing description is only an overview of the present application, and is intended to be implemented in accordance with the teachings of the present application in order that the same may be more clearly understood and to make the same and other objects, features and advantages of the present application more readily apparent.
Drawings
In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly described below, and it is apparent that the drawings in the following description are only some embodiments of the present invention, and other drawings may be obtained according to these drawings without inventive effort for a person skilled in the art. In the drawings:
FIG. 1 is a flow chart of an adaptive strategy optimization method for an autonomous robot based on interactive feedback, according to an exemplary embodiment of the present invention;
fig. 2 is a schematic structural diagram of an adaptive strategy optimization system of an autonomous robot based on interactive feedback according to an exemplary embodiment of the present invention.
Detailed Description
Further advantages and effects of the present invention will become readily apparent to those skilled in the art from the disclosure herein, by referring to the accompanying drawings and the preferred embodiments. The invention may be practiced or carried out in other embodiments that depart from the specific details, and the details of the present description may be modified or varied from the spirit and scope of the present invention. It should be understood that the preferred embodiments are presented by way of illustration only and not by way of limitation.
It should be noted that the illustrations provided in the following embodiments merely illustrate the basic concept of the present invention by way of illustration, and only the components related to the present invention are shown in the drawings and are not drawn according to the number, shape and size of the components in actual implementation, and the form, number and proportion of the components in actual implementation may be arbitrarily changed, and the layout of the components may be more complicated.
In the following description, numerous details are set forth in order to provide a more thorough explanation of embodiments of the present invention, it will be apparent, however, to one skilled in the art that embodiments of the present invention may be practiced without these specific details, in other embodiments, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the embodiments of the present invention.
Embodiment one:
the self-adaptive strategy optimization method of the robot with body based on the interaction feedback, as shown in fig. 1, comprises the following steps:
Acquiring image data, depth data, force sense data and pose data of the robot body in the task execution process, and generating environment state information, wherein the environment state information comprises the current spatial pose, perception characteristics and stress state of the robot body;
Analyzing and structuring the environment state information, a voice command, a gesture command and a binary evaluation signal from a user to generate a standardized feedback representation, and expressing the joint characteristics of the environment state and the user interaction intention;
calculating a reward function according to the environmental state information and the standardized feedback characterization, generating a standardized reward value, and quantifying the consistency relation between the environmental state information and the standardized feedback characterization;
Performing high-order mapping on the environmental state information and the standardized reward value to generate a state-feedback joint representation, and reflecting the interaction relation of the state constraint and the feedback constraint in the strategy generation process;
Generating a candidate strategy set based on the state-feedback joint representation, and screening and generating updated strategy parameters from the candidate strategy set under constraint conditions, wherein the constraint conditions comprise strategy consistency constraint and execution stability constraint;
Mapping the updated strategy parameters into a robot action sequence, and performing operation and environment interaction by the robot according to the action sequence to generate a feedback execution result;
Comparing an execution result after the robot executes the action sequence with the standardized feedback characterization to generate a consistency index, judging whether strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, generating a closed-loop data stream by the execution result and the consistency index, and supporting self-adaptive optimization of strategies.
The invention is further configured that the collecting data of the robot with body in the task execution process, and generating the environmental status information includes:
In the task execution process, collecting image data, depth data, force sense data and pose data of the robot, preprocessing the collected data to generate a unified data structure, and specifically, setting the robot to perform time steps The data of each mode is: , wherein, Is an image matrix, is acquired by a robot camera, and has the size ofRepresenting the spatial visual information, and the information of the spatial visual information,Is a depth data array, the size isDistance information indicating each point of the robot and the environment,For force sense data sequence, the length isIndicating the stress state of each joint or end effector of the robot,Is a pose matrix with the size ofWhereinFor the number of joints, 6 represents the spatial position and orientation information of each joint, and the data are subjected to unified preprocessing to ensure that the input formats are unified and avoid the influence of the cross-modal difference on the subsequent mapping;
extracting multi-scale features of image data to generate visual perception feature representation, in particular to an image matrix And (3) multi-scale convolution feature extraction: , wherein, For the number of layers of the convolution scale,Is the firstLayer convolution operations, including convolution kernels, nonlinear activation and normalization operations,Is the firstThe layer convolves the weight matrix with the same,Representing element-wise multiplication, for fusing features between convolutions,Representing a matrix for image features, the operation extracting local edge, global texture and shape information at multiple scales to form a visual feature matrix;
Gradient encoding and neighborhood feature mapping are carried out on depth data to generate space geometrical feature representation, and the depth data is specificConversion to high dimensional spatial features: , wherein, For the depth gradient operator, depth variation information is calculated,For the depth local neighborhood weighting function, depth neighborhood information is encoded,For a nonlinear mapping function, two-dimensional depth data is embedded into a high-dimensional feature space,Capturing the geometric information of the environment space through gradient operators and neighborhood coding for representing the depth characteristics, and converting the two-dimensional depth information into high-dimensional characteristics capable of participating in state characterization;
The force sense data is subjected to joint force weighting and nonlinear mapping to generate force sense state representation, and specifically, the force sense sequence is subjected to force sense Constructing a force sense state representation: , wherein, For the number of force sensing sensors,The weight coefficient is used for representing the force sense influence of different joints,For the force sense nonlinear mapping function, the original force sense numerical value is mapped to the feature space,For force sense characteristic representation, weighting nonlinear mapping is carried out on stress information of each joint or actuator to form a unified force sense characteristic matrix which reflects the stress state of the robot;
position coding and direction coding are carried out on pose data to generate space pose representation, and in particular, pose matrix A spatial mapping is performed and the spatial mapping is performed,, wherein,For joint position coding functions, spatial pose information is captured,For the joint orientation encoding function, spatial orientation information is captured,The representation is fused element by element,The position and orientation codes are fused to obtain complete space gesture features for the gesture feature matrix, and a gesture base is provided for the environment state information;
The visual perception feature representation, the space geometrical feature representation, the force sense state representation and the space posture representation are subjected to unified mapping and high-order fusion to generate environment state information, wherein the environment state information comprises the current space posture, the perception feature and the stress state of the robot, and in particular, the multi-mode feature is obtained And performing unified mapping to generate environment state information: , wherein, The characteristic stitching operation is represented as such,For high-order nonlinear mapping operators, multi-modal features are embedded into a unified state space,The system is environment state information and comprises space gestures, perception features and stress states, vision, depth, force sense and pose information are unified into high-dimensional environment state representation through feature splicing and high-order mapping, and complete environment state information of coverage space, stress and perception is formed through multi-mode fusion of vision, depth, force sense and pose.
The invention is further arranged such that the generating a normalized feedback representation comprises:
analyzing and structuring the environment state information, the voice command, the gesture command and the binary evaluation signal of the user, particularly, in the time step The environmental status information of (a) isThe user interaction signal is: , wherein, Is a voice instruction vector sequence with the length of,The feature matrix is a gesture instruction feature matrix with the size of,In order to be of a height, the height,In the form of a width, the width,In order to provide the number of channels,For binary evaluation of signal sequences, the length is;
Extracting time sequence features and non-linear coding to phonetic instruction to generate phonetic feature representation, and specifically, phonetic vector sequenceThe speech features are generated by a non-linear embedding mapping,, wherein,Is the firstFeature transformation functions of the individual speech vectors, extracting timing semantic information,For a nonlinear attention weighting function, weights are assigned according to speech importance,Representing an element-by-element multiplication,Capturing key information related to task execution in a voice instruction through element-by-element nonlinear weighting for voice characteristic representation;
mapping the space and channel characteristics of the gesture command to generate a gesture characteristic representation, and specifically, mapping the gesture matrix Performing space one-channel convolution mapping: , wherein, Extracting motion direction and amplitude characteristics for the gesture local nonlinear mapping function,For the convolution weight matrix, the importance of different spatial locations and channels is adjusted,For gesture feature representation, capturing key information of gestures on a space and a channel through space-channel weighted mapping, and realizing representation of associated features of the gestures and environmental states;
Performing nonlinear embedding and weighting processing on the binary evaluation signals to generate evaluation characteristic representations, specifically, performing nonlinear embedding and weighting processing on binary evaluation sequences Constructing a nonlinear response map: , wherein, Is a non-linear embedded function of binary signal, willOr (b)The evaluation is mapped to a high-dimensional feature space,As the weight coefficient, the importance of each evaluation time is reflected,For the binary feedback feature, mapping the simple binary evaluation into a high-dimensional feature so as to be fused with the voice and gesture feature, and reflecting the guiding function of the evaluation on policy optimization;
Performing high-order nonlinear joint mapping on the voice characteristic representation, the gesture characteristic representation, the evaluation characteristic representation and the environment state information to generate a standardized feedback representation, expressing joint characteristics of the environment state and the user interaction intention, and particularly, the environment state information High-order nonlinear joint mapping is carried out on the voice, the gesture and the binary evaluation characteristics: , wherein, The characteristic stitching operation is represented as such,Representing element-by-element interaction mapping, preserving the coupling relationship between state and feedback,For high-order nonlinear mapping operators, multi-modal features are embedded into a unified feature space,For standardized feedback characterization, the joint characteristics of the environment state and the user interaction intention are reflected, and the environment state and the multi-mode user interaction characteristics are coupled through element-by-element interaction and high-order mapping to form the joint characterization which can directly participate in policy optimization.
The invention is further arranged such that the generating a normalized prize value comprises:
the method comprises the steps of carrying out nonlinear mapping and high-order interaction on environment state information and standardized feedback characterization to form a state-feedback coupling characteristic, specifically carrying out element-by-element interaction and high-order mapping on the environment state and the feedback characteristic to generate the state-feedback coupling characteristic: , wherein, The key space gesture, the perception characteristic and the stress state are enhanced for the nonlinear mapping function of the environment state,To feedback the key information characterizing the nonlinear mapping function, enhance the user's intent,The characteristic stitching operation is represented as such,Representing element-by-element interactions, capturing the coupling relationship between state and feedback,Output state-feedback coupling feature for high order nonlinear mapping operatorEmbedding the combined information of the environment state and the user feedback into a unified feature space through element-by-element interaction and high-order mapping, quantifying the consistency potential, realizing the deep fusion of the environment state and the user feedback through state-feedback high-order coupling mapping, and ensuring that the reward function is highly sensitive to key features;
nonlinear high-order integration is performed on the state-feedback coupling characteristic to generate a reward value, and the state-feedback coupling characteristic is specifically performed Generating a prize value by performing a nonlinear high order integration: , wherein, Is the firstDimension of environmental status, 1State at the feedback feature dimension-a feedback coupling feature,For the number of dimensions of the environmental state,For the feedback of the number of feature dimensions,For a nonlinear index, control the higher order sensitivity,The importance of a feedback combination of different states is reflected as a weight matrix,Mapping the result of the integration to a prize value as a non-linear mapping functionThe multi-dimensional information is converted into a single rewarding index through high-order integral and weighted mapping of the state-feedback coupling characteristic, consistency between the environment state and user feedback is captured, and a nonlinear coupling relation is extracted from a multi-dimensional state-feedback space through high-order consistency integral, so that the problem that complex interaction information is lost by a simple weighting method is avoided;
normalizing the generated reward value to generate a normalized reward value, specifically, generating the reward value Mapping to a standardized interval to form a rewarding index capable of directly participating in policy optimization: , wherein, In order to normalize the mapping operator,Is small constant, prevents the abnormality of zero removal,The method is used for standardized reward values for subsequent state-feedback joint mapping and strategy generation, the higher-order reward values are mapped to a unified scale, stability and comparability of the reward values in the optimization process are guaranteed, the reward values have the unified numerical scale through normalized reward construction, and the reward values can be stably applied to strategy optimization under different tasks and different feedback conditions.
The invention is further arranged for generating a state-feedback joint representation comprising:
The method comprises the steps of carrying out nonlinear transformation on environment state information, strengthening state constraint characteristics, specifically, carrying out high-order nonlinear mapping on the environment state information, and highlighting the state constraint characteristics: , wherein, For the number of dimensions of the environmental state,For the number of state sub-feature dimensions,The importance of different state dimensions and sub-features is reflected for the weight coefficient matrix,Is a nonlinear exponential matrix, enhances the high-order sensitivity,For the mapped environment state characteristic representation, constraint information of the environment state in strategy generation is strengthened through exponential weighting and high-order nonlinear mapping;
non-linear mapping is carried out on the standardized rewards value, the feedback constraint characteristic is enhanced, and specifically, the standardized rewards are obtained High order nonlinear mapping is performed to emphasize feedback constraint features: , wherein, In order to map the number of terms,To weight the coefficients, reflect the contributions of the different higher-order terms to the joint representation,Is nonlinear index, controls the sensitivity of high-order feedback,Converting the standardized rewards into high-order feedback constraint features through nonlinear mapping for mapped rewards features, and coupling the high-order feedback constraint features with state features in a joint representation;
High-order interaction is carried out on the mapped environmental state characteristics and the mapped standardized rewards value to form state-feedback combined characteristics, and the environmental state characteristics are specifically With rewards featurePerforming element-by-element high-order interaction and nonlinear mapping: , wherein, The representation is an element-by-element interaction,For the bonus feature to be a non-linear spread function,The weight tensor is combined for the state-one prize,As a non-linear exponential tensor,For the state-feedback joint representation, the state constraint and the feedback constraint are fully coupled through element-by-element interaction and high-order nonlinear mapping, so that the joint representation is realized, and high-order constraint information is provided for strategy generation.
The invention further provides that the filtering and generating updated strategy parameters from the candidate strategy set under the constraint condition comprises the following steps:
Generating a candidate strategy set based on the state-feedback joint representation, and specifically, constructing a multidimensional strategy generation operator based on the state-feedback joint representation: , wherein, The number of dimensions that are jointly represented for the state-feedback,For the number of joint feature sub-dimensions,Generating a weight matrix for the strategies, adjusting the contribution of the joint features to different candidate strategies,For a nonlinear exponential matrix, control the higher order sensitivity,Random perturbation parameters generated for the strategy, for increasing candidate strategy diversity,Is the firstCandidate strategies, definitionGenerating diversified candidate strategies for the candidate strategy set through the high-order nonlinear combined joint representation features, and ensuring that the strategy set covers a potential optimal solution space;
carrying out policy consistency constraint evaluation on each candidate policy, and quantifying the consistency of the candidate policies in the continuous execution process: , wherein, For the dimension of the policy action,In order to be a consistent weight matrix,Is a nonlinear index matrix, enhances the sensitivity of high-order differences,For the strategy consistency index, quantifying the consistency of the strategy in the continuous action sequence, and ensuring that the generated strategy cannot cause action jump or instability;
Performing execution stability constraint evaluation on each candidate strategy, and quantifying the robustness of the candidate strategy in a dynamic environment: , wherein, As a dimension of the action,As a factor of the weight of the stability,Is a nonlinear index, reflects the sensitivity of high-order disturbance,Is an environment disturbance simulation value,To perform the stability index, quantifying the robustness of the strategy in a dynamic environment;
Screening policies meeting the policy consistency constraint and the execution stability constraint from the candidate policy set, generating updated policy parameters, specifically, screening policies meeting the constraint conditions from the candidate policy set, and generating updated policy parameters: , wherein, And is also provided withTo meet the policy index set of consistency and stability thresholds,To filter the policy weights, reflect the policy contribution,Is nonlinear index, enhances the high-order strategy characteristic,And generating final strategy parameters for the updated strategy parameters by high-order weighted combination of candidate strategies, and ensuring that the strategies are optimized under the constraint of consistency and stability.
The invention further provides that the generating feedback execution result comprises:
The method comprises the steps of converting updated strategy parameters into a robot action sequence, and enabling a robot to execute operation and environment interaction according to the action sequence, specifically, constructing a high-order action inference operator based on the updated strategy parameters to generate the action sequence: , wherein, The dimensions of the policy parameters and the sub-dimensions are respectively,Generating a weight matrix for the action, controlling the contribution of the policy parameters to the action generation,Is a nonlinear exponential matrix, enhances the high-order sensitivity,Generating random disturbance parameters for the motion, for increasing the motion diversity,For the environmental adaptation correction, reflecting the environmental impact of the action during execution,Is the robot NoIndividual actions, generating a complete sequence of actionsGenerating an action sequence through high-order nonlinear combination strategy parameters, and simultaneously adding random disturbance and environmental correction to realize action diversity and environmental adaptability;
in the process of executing the robot action, the environment state changes along with the higher-order influence of the action sequence to generate the environment state after the action is executed, specifically, the robot follows the action sequence Operating the environment, and updating the environment state: , wherein, The function is updated for the state of the environment,For the higher order impact operator of the action on the environmental state, reflecting the nonlinear change of the action execution on the environmental state,For the environment state after the action is executed, applying high-order nonlinear influence on the environment through the action sequence to form environment response associated with the action sequence;
generating a feedback execution result according to the environment state after the action is executed and the action sequence, quantifying the relation between the action and the environment response, and specifically generating the feedback execution result according to the environment state after the action is executed: , wherein, For feeding back the weight matrix, the importance of the corresponding relation between the environment state and the action is reflected,Is nonlinear index, enhances the sensitivity of high-order difference,And the feedback execution result is formed by quantifying the difference between the action execution and the environment response through a high-order nonlinear function, so as to provide a data basis for closed loop self-adaption.
The method further provides that the determining whether the policy parameter accords with the feedback constraint according to the consistency index and the preset threshold value comprises the following steps:
The method comprises the steps of carrying out high-order difference calculation on a feedback execution result generated after the robot executes an action sequence and a standardized feedback representation, quantifying nonlinear deviation between the action sequence and user interaction intention, and specifically carrying out nonlinear high-order difference quantification on the feedback execution result and the standardized feedback representation: , wherein, The dimensions of the action are represented and,Representing the dimension of the state of the environment,Is a nonlinear index matrix, enhances the sensitivity of high-order differences,As a weight matrix, reflects the importance of different actions and environmental dimensions,Is the firstThe action dimension is at the firstCapturing a small deviation and a complex interaction effect through the high-order difference in the individual environment dimensions and the difference between the execution result and the user expectation by the high-order nonlinear quantitative feedback;
Generating a consistency index according to the calculation result of the high-order difference, and quantifying the consistency degree of strategy parameters and feedback constraint Performing multidimensional nonlinear synthesis to generate a consistency index: , wherein, For the number of dimensions of the motion,For the number of dimensions of the environmental state,Is a second-order nonlinear exponential matrix, enhances the sensitivity to high-order coupling,Adjusting the contribution of each dimension to the consistency index for the weight matrix,For quantifying a high-order index of strategy parameter and feedback constraint consistency, the multidimensional nonlinear combination ensures that the consistency index fully reflects a complex interaction relation between an action sequence and a user expectation;
judging whether the strategy parameters meet feedback constraint according to the consistency index and a preset threshold value to form a strategy parameter closed-loop judgment result: , wherein, In order to meet the policy parameters of the feedback constraint,In order to trigger parameters for policy adjustment or optimization,For consistency judgment threshold, strategy parameter closed-loop judgment is realized by comparing the high-order consistency index with the threshold, so that the action sequence is ensured to meet the user intention and the environmental requirement;
combining the feedback execution result with the consistency index to generate a closed-loop data stream, and specifically, combining the feedback execution result with the consistency index to form the closed-loop data stream: , wherein, Is a nonlinear index matrix, enhances the sensitivity of closed-loop data to higher-order consistency,The method is a closed-loop data stream, and is used for the next round of strategy self-adaptive optimization, and the closed-loop data stream reflects the relation between strategy execution and user expectations in a high-order mode.
The invention is further arranged that the generating of the closed-loop data stream, the adaptive optimization of the support strategy comprises:
according to the closed-loop data stream, high-order error extraction is carried out on a feedback execution result, a consistency index and a strategy parameter judgment result after the robot executes an action sequence; specifically, policy bias information is extracted from a closed-loop data stream to form an error matrix: , wherein, As a dimension of the action,As a dimension of the environment,Is the first in the closed loop data streamThe action dimension is at the firstThe integrated feedback information in the individual environmental dimensions,As a function of the current policy parameters,Is a nonlinear index matrix, enhances sensitivity to higher order deviations,For the weight matrix, the contributions of different actions and environmental dimensions in the optimization are adjusted,The optimization sensitivity and precision are ensured by amplifying the small deviation in a high-order nonlinear way for the quantization strategy parameter deviating from the high-order error of the closed-loop data;
Generating nonlinear optimization increment according to the high-order error, performing self-adaptive adjustment on strategy parameters to form an updated strategy parameter set, and specifically, calculating the nonlinear optimization increment according to the high-order error: , wherein, Taking into account the higher-order dependency between motion and environmental dimensions as a nonlinear coupling function,Controlling the adjustment amplitude, preventing overshoot or oscillation,For the update amplitude of the action parameters in the round of optimization, nonlinear coupling ensures that adjustment not only depends on the dimension error, but also can respond to feedback of other dimensions, and global self-adaptive optimization is realized;
Generating a next round of robot action sequence according to the updated strategy parameters, realizing the consistency of the action sequence with the user interaction intention and the environment constraint, and specifically generating the next round of strategy parameters through closed-loop data flow and optimized increment: When the policy parameters are determined to be valid At the time of incrementThe adjustment amplitude is smaller, the stability is kept, and when the strategy parameters are judged to be invalidAt the time of incrementThe adjustment amplitude is increased, the parameters deviating from the feedback constraint are corrected rapidly, and the strategy parameters are optimized in a continuous self-adaptive mode through closed-loop error quantization and nonlinear increment adjustment, so that the action sequence is ensured to approach the user interaction intention gradually.
Embodiment two:
Referring to fig. 2, the exemplary adaptive strategy optimization system for an autonomous robot based on interactive feedback includes:
The multi-mode sensing data acquisition module acquires image data, depth data, force sense data and pose data of the robot body in the task execution process, and generates environment state information, wherein the environment state information comprises the current spatial pose, sensing characteristics and stress state of the robot body;
The interactive feedback analysis module analyzes and constructs the environment state information, the voice command, the gesture command and the binary evaluation signal from the user to generate a standardized feedback representation, and expresses the joint characteristics of the environment state and the interactive intention of the user;
the instant rewarding modeling module calculates a rewarding function according to the environment state information and the standardized feedback characterization, generates a standardized rewarding value and quantifies the consistency relation between the environment state information and the standardized feedback characterization;
The state-feedback joint mapping module is used for carrying out high-order mapping on the environmental state information and the standardized rewarding value to generate a state-feedback joint representation, and reflecting the interaction relation of the state constraint and the feedback constraint in the strategy generation process;
The constraint optimization strategy generation module generates a candidate strategy set based on the state-feedback joint representation, screens and generates updated strategy parameters from the candidate strategy set under constraint conditions, wherein the constraint conditions comprise strategy consistency constraint and execution stability constraint;
the action reasoning and control mapping module maps the updated strategy parameters into a robot action sequence, and the robot performs operation and environment interaction according to the action sequence to generate a feedback execution result;
And the execution feedback consistency verification module compares an execution result after the robot executes the action sequence with the standardized feedback characterization to generate a consistency index, judges whether strategy parameters accord with feedback constraint according to the consistency index and a preset threshold value, generates a closed-loop data stream according to the execution result and the consistency index, and supports self-adaptive optimization of the strategy.
It should be noted that, the self-adaptive strategy optimization system of the self-adaptive strategy based on the interactive feedback provided by the above embodiment and the self-adaptive strategy optimization method of the self-adaptive strategy based on the interactive feedback provided by the above embodiment belong to the same conception, the specific manner in which the respective modules and units perform operations has been described in detail in the method embodiments, and will not be described here again. In practical application, the self-adaptive strategy optimization system of the self-adaptive robot based on the interactive feedback provided by the embodiment can distribute the functions by different functional modules according to the needs, namely, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above, and the self-adaptive strategy optimization system is not limited in this place.
The foregoing is merely illustrative of the present application, and the present application is not limited thereto, and any person skilled in the art will readily recognize that variations or substitutions are within the scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims (6)

1.基于交互反馈的具身机器人自适应策略优化方法,其特征在于,包括:1. A method for adaptive strategy optimization of embodied robots based on interactive feedback, characterized by comprising: 采集具身机器人在任务执行过程中的图像数据、深度数据、力觉数据和位姿数据,生成环境状态信息,环境状态信息包含具身机器人当前的空间姿态、感知特征和受力状态;Collect image data, depth data, force data and pose data of the embodied robot during the task execution process, and generate environmental state information, which includes the embodied robot's current spatial posture, perception characteristics and force state; 将环境状态信息与来自用户的语音指令、手势指令和二值评价信号进行解析与结构化处理,生成标准化反馈表征,表达环境状态与用户交互意图的联合特征;其中,生成标准化反馈表征包括:将环境状态信息与用户的语音指令、手势指令和二值评价信号进行解析与结构化处理;对语音指令进行时序特征提取和非线性编码,生成语音特征表示;对手势指令进行空间和通道特征映射,生成手势特征表示;对二值评价信号进行非线性嵌入和加权处理,生成评价特征表示;将语音特征表示、手势特征表示、评价特征表示与环境状态信息进行高阶非线性联合映射,生成标准化反馈表征,表达环境状态与用户交互意图的联合特征;The process involves parsing and structuring environmental state information along with user voice commands, gesture commands, and binary evaluation signals to generate a standardized feedback representation that expresses the joint features of the environmental state and the user's interaction intent. This standardized feedback representation generation includes: parsing and structuring the environmental state information along with user voice commands, gesture commands, and binary evaluation signals; extracting temporal features and performing nonlinear encoding on the voice commands to generate voice feature representations; mapping spatial and channel features on the gesture commands to generate gesture feature representations; performing nonlinear embedding and weighting on the binary evaluation signals to generate evaluation feature representations; and performing a high-order nonlinear joint mapping of the voice feature representations, gesture feature representations, and evaluation feature representations with the environmental state information to generate a standardized feedback representation that expresses the joint features of the environmental state and the user's interaction intent. 根据环境状态信息与标准化反馈表征计算奖励函数,生成标准化奖励值,量化环境状态信息与标准化反馈表征的一致性关系;其中,生成标准化奖励值包括:对环境状态信息和标准化反馈表征进行非线性映射和进行高阶交互,形成状态—反馈耦合特征;对状态—反馈耦合特征进行非线性高阶积分,生成奖励值;将生成的奖励值进行归一化处理,生成标准化奖励值;The reward function is calculated based on environmental state information and standardized feedback representation to generate a standardized reward value, quantifying the consistency relationship between environmental state information and standardized feedback representation. Generating the standardized reward value includes: performing nonlinear mapping and high-order interaction on environmental state information and standardized feedback representation to form a state-feedback coupling feature; performing nonlinear high-order integral on the state-feedback coupling feature to generate a reward value; and normalizing the generated reward value to generate a standardized reward value. 将环境状态信息与标准化奖励值进行高阶映射,生成状态—反馈联合表示,反映在策略生成过程中体现状态约束与反馈约束的交互关系;By performing a high-order mapping between environmental state information and standardized reward values, a joint state-feedback representation is generated, which reflects the interaction between state constraints and feedback constraints in the policy generation process. 基于状态—反馈联合表示,生成候选策略集合,从候选策略集合中在约束条件下筛选生成更新后的策略参数,约束条件包括策略一致性约束和执行稳定性约束;Based on the state-feedback joint representation, a set of candidate policies is generated. The updated policy parameters are then generated by filtering the candidate policies under constraints, including policy consistency constraints and execution stability constraints. 将更新后的策略参数映射为机器人动作序列,机器人按照动作序列执行操作与环境交互,生成反馈执行结果;The updated policy parameters are mapped to robot action sequences. The robot performs operations and interacts with the environment according to the action sequences, generating feedback execution results. 将机器人执行动作序列后生成的反馈执行结果与标准化反馈表征进行比对,生成一致性指标,根据一致性指标和预设阈值判定策略参数是否符合反馈约束,其中,根据一致性指标和预设阈值判定策略参数是否符合反馈约束包括:将机器人执行动作序列后生成的反馈执行结果与标准化反馈表征进行高阶差异计算,量化动作序列与用户交互意图之间的非线性偏离;根据高阶差异计算结果生成一致性指标,量化策略参数与反馈约束的一致性程度;根据一致性指标与预设阈值判定策略参数是否符合反馈约束,形成策略参数闭环判定结果;将反馈执行结果与一致性指标结合生成闭环数据流,反馈执行结果与一致性指标生成闭环数据流,支持策略的自适应优化,其中,生成闭环数据流,支持策略的自适应优化包括:依据闭环数据流,对机器人执行动作序列后的反馈执行结果、一致性指标和策略参数判定结果进行高阶误差提取;根据高阶误差生成非线性优化增量,对策略参数进行自适应调整,形成更新后的策略参数集合;根据更新后的策略参数生成下一轮机器人动作序列,实现动作序列与用户交互意图和环境约束的一致性。The feedback execution results generated after the robot executes a sequence of actions are compared with standardized feedback representations to generate a consistency index. Based on the consistency index and a preset threshold, it is determined whether the strategy parameters meet the feedback constraints. This determination includes: performing high-order difference calculations on the feedback execution results and standardized feedback representations to quantify the non-linear deviation between the action sequence and the user's interaction intent; generating a consistency index based on the high-order difference calculation results to quantify the degree of consistency between the strategy parameters and the feedback constraints; and determining whether the strategy parameters meet the feedback constraints based on the consistency index and the preset threshold, thus forming a strategy... The closed-loop determination result of the parameters is omitted; the feedback execution result and the consistency index are combined to generate a closed-loop data stream, which supports the adaptive optimization of the strategy. The generation of the closed-loop data stream and the support for the adaptive optimization of the strategy include: extracting higher-order errors from the feedback execution result, consistency index and strategy parameter determination result after the robot executes the action sequence based on the closed-loop data stream; generating nonlinear optimization increments based on the higher-order errors to adaptively adjust the strategy parameters and form an updated set of strategy parameters; generating the next round of robot action sequence based on the updated strategy parameters to achieve consistency between the action sequence and the user's interaction intention and environmental constraints. 2.根据权利要求1所述的基于交互反馈的具身机器人自适应策略优化方法,其特征在于,采集具身机器人在任务执行过程中的数据,生成环境状态信息包括:2. The adaptive strategy optimization method for embodied robots based on interactive feedback according to claim 1, characterized in that, collecting data from the embodied robot during task execution and generating environmental state information includes: 在任务执行过程中,采集具身机器人的图像数据、深度数据、力觉数据和位姿数据,对采集的数据进行预处理,生成统一的数据结构;During the task execution, image data, depth data, force data and pose data of the embodied robot are collected, and the collected data are preprocessed to generate a unified data structure. 对图像数据进行多尺度特征提取,生成视觉感知特征表示;Multi-scale feature extraction is performed on image data to generate visual perception feature representations; 对深度数据进行梯度编码与邻域特征映射,生成空间几何特征表示;Gradient encoding and neighborhood feature mapping are performed on depth data to generate spatial geometric feature representations; 对力觉数据进行关节受力加权与非线性映射,生成力觉状态表示;Force perception data is weighted at joints and nonlinearly mapped to generate a force perception state representation; 对位姿数据进行位置编码与方向编码,生成空间姿态表示;Position and orientation encoding are performed on the pose data to generate a spatial pose representation; 将视觉感知特征表示、空间几何特征表示、力觉状态表示和空间姿态表示进行统一映射和高阶融合,生成环境状态信息,环境状态信息包含具身机器人当前的空间姿态、感知特征和受力状态。Visual perception feature representation, spatial geometric feature representation, force state representation and spatial posture representation are uniformly mapped and fused in a higher order to generate environmental state information, which includes the embodied robot's current spatial posture, perception features and force state. 3.根据权利要求1所述的基于交互反馈的具身机器人自适应策略优化方法,其特征在于,生成状态—反馈联合表示包括:3. The adaptive strategy optimization method for embodied robots based on interactive feedback according to claim 1, characterized in that generating the state-feedback joint representation includes: 对环境状态信息进行非线性变换,强化状态约束特征;Nonlinear transformations are applied to environmental state information to enhance state constraint characteristics; 对标准化奖励值进行非线性映射,增强反馈约束特征;A nonlinear mapping is applied to the standardized reward value to enhance the feedback constraint characteristics; 将映射后的环境状态特征与映射后的标准化奖励值进行高阶交互,形成状态—反馈联合特征。The mapped environmental state features are interacted with the mapped standardized reward value in a higher order to form a state-feedback joint feature. 4.根据权利要求3所述的基于交互反馈的具身机器人自适应策略优化方法,其特征在于,从候选策略集合中在约束条件下筛选生成更新后的策略参数包括:4. The adaptive strategy optimization method for embodied robots based on interactive feedback according to claim 3, characterized in that, the process of selecting and generating updated strategy parameters from the candidate strategy set under constraints includes: 基于状态—反馈联合表示生成候选策略集合;Generate a candidate policy set based on state-feedback joint representation; 对每个候选策略进行策略一致性约束评估,量化候选策略在连续执行过程中的连贯性;For each candidate strategy, a strategy consistency constraint evaluation is performed to quantify the consistency of the candidate strategy during continuous execution. 对每个候选策略进行执行稳定性约束评估,量化候选策略在动态环境中的鲁棒性;For each candidate strategy, an execution stability constraint evaluation is performed to quantify the robustness of the candidate strategy in a dynamic environment; 从候选策略集合中筛选满足策略一致性约束和执行稳定性约束的策略,生成更新后的策略参数。Select strategies from the candidate strategy set that satisfy the strategy consistency constraint and the execution stability constraint, and generate updated strategy parameters. 5.根据权利要求1所述的基于交互反馈的具身机器人自适应策略优化方法,其特征在于,生成反馈执行结果包括:5. The adaptive strategy optimization method for embodied robots based on interactive feedback according to claim 1, characterized in that generating feedback execution results includes: 将更新后的策略参数转化为机器人动作序列,机器人按照动作序列执行操作与环境交互;The updated policy parameters are converted into robot action sequences, and the robot performs operations and interacts with the environment according to the action sequences. 机器人动作执行过程中,环境状态随动作序列的高阶影响而变化,生成动作执行后的环境状态;During the execution of robot actions, the environmental state changes with the higher-order influences of the action sequence, generating the environmental state after the action is executed; 根据动作执行后的环境状态与动作序列生成反馈执行结果,量化动作与环境响应之间的关系。Feedback execution results are generated based on the environmental state and action sequence after the action is executed, quantifying the relationship between the action and the environmental response. 6.基于交互反馈的具身机器人自适应策略优化系统,用于实现权利要求1-5任一项所述的基于交互反馈的具身机器人自适应策略优化方法,其特征在于,包括:6. An embodied robot adaptive strategy optimization system based on interactive feedback, used to implement the embodied robot adaptive strategy optimization method based on interactive feedback as described in any one of claims 1-5, characterized in that it comprises: 多模态感知数据采集模块,采集具身机器人在任务执行过程中的图像数据、深度数据、力觉数据和位姿数据,生成环境状态信息,环境状态信息包含具身机器人当前的空间姿态、感知特征和受力状态;The multimodal perception data acquisition module collects image data, depth data, force data, and pose data of the embodied robot during task execution, and generates environmental state information, which includes the embodied robot's current spatial posture, perception features, and force state. 交互反馈解析模块,将环境状态信息与来自用户的语音指令、手势指令和二值评价信号进行解析与结构化处理,生成标准化反馈表征,表达环境状态与用户交互意图的联合特征;The interactive feedback parsing module parses and structures environmental state information along with voice commands, gesture commands, and binary evaluation signals from the user to generate standardized feedback representations that express the joint features of environmental state and user interaction intent. 即时奖励建模模块,根据环境状态信息与标准化反馈表征计算奖励函数,生成标准化奖励值,量化环境状态信息与标准化反馈表征的一致性关系;The instant reward modeling module calculates the reward function based on environmental state information and standardized feedback representation, generates standardized reward values, and quantifies the consistency relationship between environmental state information and standardized feedback representation. 状态—反馈联合映射模块,将环境状态信息与标准化奖励值进行高阶映射,生成状态—反馈联合表示,反映在策略生成过程中体现状态约束与反馈约束的交互关系;The state-feedback joint mapping module performs a high-order mapping between environmental state information and standardized reward values to generate a state-feedback joint representation, which reflects the interaction between state constraints and feedback constraints in the policy generation process. 约束优化策略生成模块,基于状态—反馈联合表示,生成候选策略集合,从候选策略集合中在约束条件下筛选生成更新后的策略参数,约束条件包括策略一致性约束和执行稳定性约束;The constraint optimization strategy generation module generates a set of candidate strategies based on the state-feedback joint representation. It then selects updated strategy parameters from the set of candidate strategies under constraints, including strategy consistency constraints and execution stability constraints. 动作推理与控制映射模块,将更新后的策略参数映射为机器人动作序列,机器人按照动作序列执行操作与环境交互,生成反馈执行结果;The action reasoning and control mapping module maps the updated policy parameters to robot action sequences. The robot performs operations and interacts with the environment according to the action sequences, generating feedback execution results. 执行反馈一致性验证模块,将机器人执行动作序列后的执行结果与标准化反馈表征进行比对,生成一致性指标,根据一致性指标和预设阈值判定策略参数是否符合反馈约束,执行结果与一致性指标生成闭环数据流,支持策略的自适应优化。The execution feedback consistency verification module compares the execution results of the robot's action sequence with the standardized feedback representation to generate a consistency index. Based on the consistency index and a preset threshold, it determines whether the strategy parameters meet the feedback constraints. The execution results and the consistency index generate a closed-loop data stream, supporting adaptive optimization of the strategy.
CN202511533027.1A 2025-10-24 2025-10-24 Self-adaptive strategy optimization method and system for body robot based on interactive feedback Active CN121223793B (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
CN202511533027.1A CN121223793B (en) 2025-10-24 2025-10-24 Self-adaptive strategy optimization method and system for body robot based on interactive feedback

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
CN202511533027.1A CN121223793B (en) 2025-10-24 2025-10-24 Self-adaptive strategy optimization method and system for body robot based on interactive feedback

Publications (2)

Publication Number Publication Date
CN121223793A CN121223793A (en) 2025-12-30
CN121223793B true CN121223793B (en) 2026-04-10

Family

ID=98159981

Family Applications (1)

Application Number Title Priority Date Filing Date
CN202511533027.1A Active CN121223793B (en) 2025-10-24 2025-10-24 Self-adaptive strategy optimization method and system for body robot based on interactive feedback

Country Status (1)

Country Link
CN (1) CN121223793B (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114995657A (en) * 2022-07-18 2022-09-02 湖南大学 A multi-modal fusion natural interaction method, system and medium for intelligent robot
CN119260754A (en) * 2024-09-18 2025-01-07 台州爱鑫智能科技有限公司 An information interaction method for embodied intelligent robots

Family Cites Families (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9274607B2 (en) * 2013-03-15 2016-03-01 Bruno Delean Authenticating a user using hand gesture
EP4017689B1 (en) * 2019-09-30 2024-10-30 Siemens Aktiengesellschaft Robotics control system and method for training said robotics control system
CN117549310A (en) * 2023-12-28 2024-02-13 亿嘉和科技股份有限公司 General system of intelligent robot with body, construction method and use method
US20250292097A1 (en) * 2024-03-14 2025-09-18 International Business Machines Corporation Optimizing grayscale release strategies based on multiple objectives and constraints
CN118567477A (en) * 2024-05-31 2024-08-30 广东铂锶特科技有限公司 Adaptive augmented reality diversified interaction method and system for complex scenes
CN120095830B (en) * 2025-04-22 2025-12-23 节卡未来科技(上海)有限公司 Robot control method, computer device, and computer-readable storage medium
CN120630670A (en) * 2025-05-07 2025-09-12 南京康龙威科技实业有限公司 A robot gait training method and system based on reinforcement learning
CN120542396A (en) * 2025-05-20 2025-08-26 平安科技(深圳)有限公司 Natural language driven table processing method, device, equipment and medium
CN120560339A (en) * 2025-05-27 2025-08-29 上海大学 A UAV swarm collaborative coverage method with dynamic obstacle avoidance
CN120807320A (en) * 2025-05-30 2025-10-17 郑州大学第五附属医院 Low-count PET image quality enhancement method and system based on fusion multi-input cycle consistency generation countermeasure network
CN120773064B (en) * 2025-09-02 2026-04-10 深圳乾海格致科技有限公司 Self-adaptive behavior mode adjusting system of intelligent robot with body

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114995657A (en) * 2022-07-18 2022-09-02 湖南大学 A multi-modal fusion natural interaction method, system and medium for intelligent robot
CN119260754A (en) * 2024-09-18 2025-01-07 台州爱鑫智能科技有限公司 An information interaction method for embodied intelligent robots

Also Published As

Publication number Publication date
CN121223793A (en) 2025-12-30

Similar Documents

Publication Publication Date Title
Hewing et al. Learning-based model predictive control: Toward safe learning in control
CN111433689B (en) Generation of control systems for target systems
KR102239186B1 (en) System and method for automatic control of robot manipulator based on artificial intelligence
CN117709185A (en) Digital twin architecture design method for process industry
EP3614314B1 (en) Method and apparatus for generating chemical structure using neural network
CN120688015A (en) AI interactive data processing system based on multimodal perception and dynamic decision-making
CN120552071A (en) Real-time collaborative decision-making method for humanoid robots based on multimodal perception fusion
CN116502069B (en) Haptic time sequence signal identification method based on deep learning
JP2020155010A (en) Neural network model contraction device
CN119974028B (en) Self-adaptive gravity center adjusting method and system for transfer robot
CN114091554A (en) Training set processing method and device
CN119442144A (en) A reinforcement learning modeling method and device integrating multi-source heterogeneous data
CN117875407A (en) A multimodal continuous learning method, device, equipment and storage medium
CN119681909A (en) Robot control method based on multi-modal large model
CN118940220A (en) A multimodal industrial data fusion method and system for discrete manufacturing
Wu et al. A framework of improving human demonstration efficiency for goal-directed robot skill learning
CN121223793B (en) Self-adaptive strategy optimization method and system for body robot based on interactive feedback
CN120388565B (en) Voice interaction method and system based on 3D (three-dimensional) virtual
CN119407779A (en) Flexible robot control method and system based on multimodal data and heuristic graph search
JP7623619B2 (en) Learning device, estimation device, learning method, estimation method, and program
Grimble et al. Non-linear predictive control for manufacturing and robotic applications
Tang Research on Computer-Aided Architectural Design Optimization Based on Building Information Modeling (BIM) Technology
CN121267944B (en) Robot motion control strategy network training method and device based on state representation learning
CN121515213A (en) An Adaptive Control Method and System for Industrial Robots Based on Multimodal Sensor Fusion
Iannotti Development and Assessment of Structural Theories Using Machine Learning Techniques

Legal Events

Date Code Title Description
PB01 Publication
PB01 Publication
SE01 Entry into force of request for substantive examination
SE01 Entry into force of request for substantive examination
GR01 Patent grant
GR01 Patent grant