Deprecated: The each() function is deprecated. This message will be suppressed on further calls in /home/zhenxiangba/zhenxiangba.com/public_html/phproxy-improved-master/index.php on line 456
CN120807875A - Target detection method and device based on color point cloud, electronic equipment and medium - Google Patents
[go: Go Back, main page]

CN120807875A - Target detection method and device based on color point cloud, electronic equipment and medium - Google Patents

Target detection method and device based on color point cloud, electronic equipment and medium

Info

Publication number
CN120807875A
CN120807875A CN202510819580.5A CN202510819580A CN120807875A CN 120807875 A CN120807875 A CN 120807875A CN 202510819580 A CN202510819580 A CN 202510819580A CN 120807875 A CN120807875 A CN 120807875A
Authority
CN
China
Prior art keywords
preset
data
point cloud
target
dimensional
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
CN202510819580.5A
Other languages
Chinese (zh)
Inventor
董小瑜
赵金璐
吕颖
刘欢
杨敬堯
刘禄
梅博然
王禹琳
陈思
杨梓凝
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
FAW Group Corp
Original Assignee
FAW Group Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by FAW Group Corp filed Critical FAW Group Corp
Priority to CN202510819580.5A priority Critical patent/CN120807875A/en
Publication of CN120807875A publication Critical patent/CN120807875A/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G01MEASURING; TESTING
    • G01SRADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
    • G01S17/00Systems using the reflection or reradiation of electromagnetic waves other than radio waves, e.g. lidar systems
    • G01S17/86Combinations of lidar systems with systems other than lidar, radar or sonar, e.g. with direction finders
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/20Image preprocessing
    • G06V10/25Determination of region of interest [ROI] or a volume of interest [VOI]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/44Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • G06V10/443Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components by matching or filtering
    • G06V10/449Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters
    • G06V10/451Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters with interaction between the filter responses, e.g. cortical complex cells
    • G06V10/454Integrating the filters into a hierarchical structure, e.g. convolutional neural networks [CNN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • G06V10/56Extraction of image or video features relating to colour
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/766Arrangements for image or video recognition or understanding using pattern recognition or machine learning using regression, e.g. by projecting features on hyperplanes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/7715Feature extraction, e.g. by transforming the feature space, e.g. multi-dimensional scaling [MDS]; Mappings, e.g. subspace methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • G06V10/806Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/50Context or environment of the image
    • G06V20/56Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
    • G06V20/58Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/60Type of objects
    • G06V20/64Three-dimensional [3D] objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/70Labelling scene content, e.g. deriving syntactic or semantic representations
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V2201/00Indexing scheme relating to image or video recognition or understanding
    • G06V2201/07Target detection

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Evolutionary Computation (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • General Health & Medical Sciences (AREA)
  • Computing Systems (AREA)
  • Software Systems (AREA)
  • Medical Informatics (AREA)
  • Databases & Information Systems (AREA)
  • Biomedical Technology (AREA)
  • Molecular Biology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Computational Linguistics (AREA)
  • Data Mining & Analysis (AREA)
  • Biophysics (AREA)
  • General Engineering & Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Radar, Positioning & Navigation (AREA)
  • Remote Sensing (AREA)
  • Biodiversity & Conservation Biology (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Electromagnetism (AREA)
  • Image Analysis (AREA)

Abstract

本申请涉及一种基于彩色点云的目标检测方法、装置、电子设备及介质。该方法包括:基于目标检测区域的激光雷达数据和RGB图像数据得到彩色点云数据;对彩色点云数据进行体素化处理,得到结构化数据,并利用预设的三维编码器网络对结构化数据进行压缩和特征提取操作,得到预设维度的特征张量;基于预设维度的特征张量生成俯视图,并基于俯视图,利用预设的中心点检测策略和预设的尺寸回归策略得到目标热力图,进而得到目标检测区域中的每个目标物体的中心点坐标信息和状态信息。由此,通过融合相机数据和激光雷达数据进行目标检测,解决了仅依赖激光雷达数据在复杂环境下目标检测精度低的问题,提高了目标检测的准确性和鲁棒性。

The present application relates to a method, device, electronic device and medium for target detection based on color point cloud. The method includes: obtaining color point cloud data based on laser radar data and RGB image data of the target detection area; voxelizing the color point cloud data to obtain structured data, and using a preset three-dimensional encoder network to compress and extract features from the structured data to obtain a feature tensor of preset dimensions; generating a top view based on the feature tensor of preset dimensions, and based on the top view, obtaining a target heat map using a preset center point detection strategy and a preset size regression strategy, thereby obtaining the center point coordinate information and status information of each target object in the target detection area. Thus, by fusing camera data and laser radar data for target detection, the problem of low target detection accuracy in complex environments when relying solely on laser radar data is solved, and the accuracy and robustness of target detection are improved.

Description

Target detection method and device based on color point cloud, electronic equipment and medium
Technical Field
The application relates to the technical field of computer vision, in particular to a target detection method, device, electronic equipment and medium based on color point cloud.
Background
The environment sensing technology of the unmanned vehicle mainly realizes detection of surrounding environment by means of external sensors such as a laser radar, a camera, a millimeter wave radar and the like, ensures that the unmanned vehicle can timely and accurately sense potential safety hazards existing in the road surface environment, quickly takes measures to avoid traffic accidents, and has irreplaceable functions for ensuring safe running of the unmanned vehicle, wherein the environment sensing is equivalent to eyes of the unmanned vehicle.
In the related art, the unmanned vehicle environment sensing can divide a 3D space into discrete voxel units based on a target detection algorithm of a laser radar, and then a deep learning model is applied to perform feature extraction and classification.
However, the method only depends on laser radar data to perform target detection, and the target detection accuracy is low in a complex environment, so that the method needs to be solved.
Disclosure of Invention
The application provides a target detection method, device, electronic equipment and medium based on color point cloud, which are used for solving the problem of low target detection precision under a complex environment only depending on laser radar data and improving the accuracy and robustness of target detection.
In order to achieve the above objective, an embodiment of a first aspect of the present application provides a target detection method based on color point cloud, including the following steps:
acquiring laser radar data and RGB image data of a target detection area, and acquiring color point cloud data based on the laser radar data and the RGB image data;
Carrying out voxelization on the color point cloud data to obtain structured data, and carrying out compression and feature extraction operation on the structured data by utilizing a preset three-dimensional encoder network to obtain a feature tensor with preset dimension;
Generating a top view based on the characteristic tensor of the preset dimension, obtaining a target thermodynamic diagram by utilizing a preset center point detection strategy and a preset size regression strategy based on the top view, and obtaining center point coordinate information and state information of each target object in the target detection area based on the target thermodynamic diagram.
By the technical means, not only the space information provided by the laser radar data is utilized, but also the color information provided by the RGB image data is combined, so that the color point cloud data can reflect the characteristics of the target object more comprehensively, and the voxelization processing converts the color point cloud data into the structured data, so that the follow-up processing is facilitated, and the accurate detection of the target object in the target detection area is realized.
According to one embodiment of the present application, the obtaining color point cloud data based on the laser radar data and RGB image data includes:
Acquiring a first internal parameter and a first external parameter of the laser radar, and acquiring a second internal parameter and a second external parameter of the camera;
Based on the first internal parameter, the first external parameter, the second internal parameter and the second external parameter, projecting the three-dimensional coordinate corresponding to the laser radar data to the two-dimensional coordinate system corresponding to the RGB image data, and based on a preset pixel-level RGB assignment strategy, assigning RGB values of corresponding pixels to each laser radar data projected to the two-dimensional coordinate system corresponding to the RGB image data, and generating the color point cloud data.
Through the technical means, the laser radar data and the RGB image data are effectively fused, and a solid foundation is provided for subsequent feature extraction and target detection.
According to an embodiment of the present application, the voxel processing is performed on the color point cloud data to obtain structured data, including:
and dispersing the three-dimensional coordinate system of the color point cloud data into a plurality of cube units which are arranged according to a preset arrangement rule based on a preset spatial resolution condition, and obtaining the structured data based on the plurality of cube units which are arranged according to the preset arrangement rule.
Through the technical means, the dimensionality of color point cloud data is effectively reduced, the data redundancy is reduced, and the data processing efficiency is improved. Meanwhile, the formation of the structured data enables the subsequent feature extraction and target detection processes to be more convenient and accurate, and powerful technical support is provided for target detection tasks in complex scenes.
According to an embodiment of the present application, the obtaining, based on the top view, a target thermodynamic diagram using a preset center point detection strategy and a preset size regression strategy includes:
identifying at least one preset key point in the top view by utilizing the preset center point strategy;
Predicting three-dimensional boundary box information of a target object corresponding to each preset key point by utilizing the preset size regression strategy based on the at least one preset key point;
And obtaining the target thermodynamic diagram by utilizing a preset dynamic Gaussian function radius strategy based on the three-dimensional boundary box information.
Through the technical means, the position of the target object in the top view can be more accurately positioned, and the three-dimensional size information of the target object is predicted. The generation mode of the target thermodynamic diagram not only improves the accuracy of target detection, but also enhances the stability and the robustness of target detection. In complex and changeable scenes, the target detection device can quickly identify a target object, and reliable data support is provided for subsequent target tracking and behavior analysis.
According to one embodiment of the present application, when obtaining the center point coordinate information and the state information of each target object in the target detection area based on the target thermodynamic diagram, the method further includes:
based on the three-dimensional boundary frame information of each target object in the target thermodynamic diagram, carrying out feature extraction operation on the three-dimensional center point of each surface of each three-dimensional boundary frame to obtain a feature value of each three-dimensional center point;
and stacking the characteristic values of each three-dimensional center point based on a bilinear interpolation strategy to obtain characteristic vectors, and inputting the characteristic vectors into a preset neural network to obtain optimized three-dimensional boundary box information of each target object.
By the technical means, the position and the size information of the target object in the three-dimensional space can be further accurate, and the optimized three-dimensional boundary box information not only improves the accuracy of target detection, but also provides a more accurate data basis for subsequent three-dimensional reconstruction and behavior analysis.
According to the target detection method based on the color point cloud, the color point cloud data can be obtained through laser radar data and RGB image data based on the target detection area, the color point cloud data are subjected to voxelization to obtain structural data, a preset three-dimensional encoder network is utilized to compress and extract features of the structural data to obtain feature tensors of preset dimensions, a top view is generated based on the feature tensors of the preset dimensions, a target thermodynamic diagram is obtained based on the top view by utilizing a preset center point detection strategy and a preset size regression strategy, and then center point coordinate information and state information of each target object in the target detection area are obtained. Therefore, the problem of low target detection precision under the complex environment only depending on the laser radar data is solved by fusing the camera data and the laser radar data to perform target detection, and the accuracy and the robustness of target detection are improved.
To achieve the above object, an embodiment of a second aspect of the present application provides a target detection device based on color point cloud, including:
The first acquisition module is used for acquiring laser radar data and RGB image data of the target detection area and acquiring color point cloud data based on the laser radar data and the RGB image data;
The processing module is used for carrying out voxelization on the color point cloud data to obtain structured data, and carrying out compression and feature extraction operation on the structured data by utilizing a preset three-dimensional encoder network to obtain a feature tensor with preset dimension;
the second obtaining module is configured to generate a top view based on the feature tensor of the preset dimension, obtain a target thermodynamic diagram based on the top view by using a preset center point detection strategy and a preset size regression strategy, and obtain center point coordinate information and state information of each target object in the target detection area based on the target thermodynamic diagram.
According to one embodiment of the present application, the first obtaining module is specifically configured to:
Acquiring a first internal parameter and a first external parameter of the laser radar, and acquiring a second internal parameter and a second external parameter of the camera;
Based on the first internal parameter, the first external parameter, the second internal parameter and the second external parameter, projecting the three-dimensional coordinate corresponding to the laser radar data to the two-dimensional coordinate system corresponding to the RGB image data, and based on a preset pixel-level RGB assignment strategy, assigning RGB values of corresponding pixels to each laser radar data projected to the two-dimensional coordinate system corresponding to the RGB image data, and generating the color point cloud data.
According to one embodiment of the application, the processing module is specifically configured to:
and dispersing the three-dimensional coordinate system of the color point cloud data into a plurality of cube units which are arranged according to a preset arrangement rule based on a preset spatial resolution condition, and obtaining the structured data based on the plurality of cube units which are arranged according to the preset arrangement rule.
According to one embodiment of the present application, the second obtaining module is specifically configured to:
identifying at least one preset key point in the top view by utilizing the preset center point strategy;
Predicting three-dimensional boundary box information of a target object corresponding to each preset key point by utilizing the preset size regression strategy based on the at least one preset key point;
And obtaining the target thermodynamic diagram by utilizing a preset dynamic Gaussian function radius strategy based on the three-dimensional boundary box information.
According to an embodiment of the present application, when obtaining the center point coordinate information and the state information of each target object in the target detection area based on the target thermodynamic diagram, the second obtaining module is further configured to:
based on the three-dimensional boundary frame information of each target object in the target thermodynamic diagram, carrying out feature extraction operation on the three-dimensional center point of each surface of each three-dimensional boundary frame to obtain a feature value of each three-dimensional center point;
and stacking the characteristic values of each three-dimensional center point based on a bilinear interpolation strategy to obtain characteristic vectors, and inputting the characteristic vectors into a preset neural network to obtain optimized three-dimensional boundary box information of each target object.
According to the target detection device based on the color point cloud, the color point cloud data can be obtained through laser radar data and RGB image data based on the target detection area, the color point cloud data are subjected to voxelization to obtain structural data, a preset three-dimensional encoder network is utilized to compress and extract features of the structural data to obtain feature tensors of preset dimensions, a top view is generated based on the feature tensors of the preset dimensions, a target thermodynamic diagram is obtained based on the top view by utilizing a preset center point detection strategy and a preset size regression strategy, and then center point coordinate information and state information of each target object in the target detection area are obtained. Therefore, the problem of low target detection precision under the complex environment only depending on the laser radar data is solved by fusing the camera data and the laser radar data to perform target detection, and the accuracy and the robustness of target detection are improved.
To achieve the above object, an embodiment of a third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor executes the program to implement the color point cloud-based target detection method according to the above embodiment.
To achieve the above object, a fourth aspect of the present application provides a computer-readable storage medium having stored thereon a computer program for being executed by a processor for implementing the color point cloud-based object detection method according to the above embodiment.
Additional aspects and advantages of the application will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the application.
Drawings
The foregoing and/or additional aspects and advantages of the application will become apparent and readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings, in which:
fig. 1 is a flowchart of a target detection method based on color point cloud according to an embodiment of the present application;
FIG. 2 is a flowchart of another object detection method based on color point cloud according to an embodiment of the present application;
fig. 3 is a schematic block diagram of a target detection device based on color point cloud according to an embodiment of the present application;
Fig. 4 is a schematic structural diagram of an electronic device according to an embodiment of the present application.
Reference numeral 10-a target detection device based on color point cloud, 100-a first acquisition module, 200-a processing module, 300-a second acquisition module, 401-a memory, 402-a processor, 403-a communication interface.
Detailed Description
Embodiments of the present application are described in detail below, examples of which are illustrated in the accompanying drawings, wherein like or similar reference numerals refer to like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the drawings are illustrative and intended to explain the present application and should not be construed as limiting the application.
The method, the device, the electronic equipment and the medium for detecting the target based on the color point cloud according to the embodiment of the application are described below with reference to the accompanying drawings, and the method for detecting the target based on the color point cloud according to the embodiment of the application is described first with reference to the accompanying drawings.
FIG. 1 is a flow chart of a method of color point cloud-based target detection in accordance with one embodiment of the present application.
Exemplary, as shown in fig. 1, the color point cloud-based target detection method includes the following steps:
in step S101, laser radar data and RGB image data of a target detection area are acquired, and color point cloud data is obtained based on the laser radar data and the RGB image data.
Specifically, in order to achieve accurate identification of the target detection area, in the embodiment of the present application, three-dimensional point cloud data (i.e., laser radar data) including spatial coordinate information of each point may be generated by scanning the target detection area with a laser radar. Meanwhile, the same target detection area may be photographed with a camera, generating RGB image data including color information of each pixel. By combining these two types of data, color point cloud data can be further processed and generated, thereby providing rich information for subsequent target detection and analysis.
It can be understood that by performing target detection by applying color point cloud, each point in the color point cloud not only contains coordinate information of three-dimensional space, but also adds RGB color information, and the point cloud data combined with the color information makes the target detection process more efficient and accurate. The application of the color point cloud target detection algorithm can bring a plurality of benefits, namely the color point cloud can effectively integrate the multi-sensor data by fusing the data from various sensors, so that the overall perception capability of the system is greatly enhanced. In complex and changeable environments such as city streets, the addition of color information can help to improve detection accuracy, and can help to filter out background noise, so that the target is more prominent. In addition, color information plays an important role in handling occlusion problems, which can help algorithms better identify and distinguish occluded objects. For those objects that are sparse in the point cloud due to a large distance or small volume, the color information provides additional visual cues that are critical to improving the detection accuracy so that the algorithm can maintain high detection performance even when facing the sparse point cloud object.
For ease of understanding, how color point cloud data is derived based on lidar data and RGB image data is described in detail below.
As one possible implementation manner, in some embodiments, color point cloud data is obtained based on laser radar data and RGB image data, and the color point cloud data is generated by acquiring a first internal parameter and a first external parameter of the laser radar and acquiring a second internal parameter and a second external parameter of a camera, projecting three-dimensional coordinates corresponding to the laser radar data to a two-dimensional coordinate system corresponding to the RGB image data based on the first internal parameter, the first external parameter, the second internal parameter and the second external parameter, and assigning RGB values of corresponding pixels to each laser radar data projected to the two-dimensional coordinate system corresponding to the RGB image data based on a preset pixel-level RGB assignment strategy.
Specifically, in the process of obtaining color point cloud data based on laser radar data and RGB image data, first, internal parameters (i.e., first internal parameters) and external parameters (i.e., first external parameters) of the laser radar device may be acquired. At the same time, the internal parameters (i.e., second internal parameters) and the external parameters (i.e., second external parameters) of the camera device may also be acquired. After the key parameters are mastered, the point cloud data in the three-dimensional space captured by the laser radar can be mapped into a two-dimensional RGB image data coordinate system through mathematical transformation, then the RGB value of each point cloud mapped onto the RGB image data corresponding to the pixel in the RGB image data is extracted, and through the process, color point cloud data can be obtained. The color point cloud data not only maintains the original three-dimensional space information, but also contains rich color information, thereby providing more visual and rich data support for subsequent image processing and analysis. If the point cloud is occluded (e.g., projected outside the image edge), the process may be by interpolation or dropping.
It will be appreciated that internal parameters are used to correct the distortion of the sensor itself, laser radar internal parameters include scan angle resolution, range accuracy, etc., camera internal parameters include focal length, principal point coordinates, distortion coefficients, etc., and external parameters are used to convert the laser radar data and RGB image data into the same coordinate system (i.e., coordinate system alignment) to align the two coordinate systems. For example, the 3D point cloud coordinate system of the lidar needs to be converted to the 2D image coordinate system of the camera.
In step S102, the color point cloud data is subjected to voxel processing to obtain structured data, and the structured data is compressed and feature extracted by using a preset three-dimensional encoder network to obtain a feature tensor of a preset dimension.
That is, after obtaining color point cloud data, the point cloud data that is originally irregularly distributed may be converted into structured data by a voxelization process. Then, a preset three-dimensional encoder network (such as the three-dimensional encoder network in the related art) can be used for compressing and extracting features of the structured data, and finally, the high-dimensional point cloud data is converted into a lower-dimensional feature representation (i.e. a feature tensor in a preset dimension) so as to facilitate subsequent processing.
Next, how to voxel the color point cloud data to obtain structured data will be described in detail.
As one possible implementation manner, in some embodiments, the voxelization is performed on the color point cloud data to obtain structured data, which includes dispersing a three-dimensional coordinate system of the color point cloud data into a plurality of cube units arranged according to a preset arrangement rule based on a preset spatial resolution condition, and obtaining the structured data based on the plurality of cube units arranged according to the preset arrangement rule.
It is understood that voxelization is a method of dividing a three-dimensional space into small cubes (voxels), similar to pixels in a two-dimensional image. Specifically, when the color point cloud data is subjected to voxel processing, the three-dimensional coordinate system where the color point cloud data is located can be subjected to discretization processing based on a preset spatial resolution condition (which can be calibrated), that is, the whole three-dimensional space is divided into a plurality of small cube units (namely voxels), the cube units are arranged according to a certain arrangement rule, and each cube unit comprises a small space region. By such processing, the otherwise continuous color point cloud data may be converted into a series of discrete cube cells, each cell containing corresponding color and location information. Finally, based on the cube units arranged according to the preset arrangement rule, structured data can be obtained, so that subsequent data processing and analysis are more convenient and efficient.
In step S103, a top view is generated based on the feature tensor of the preset dimension, a target thermodynamic diagram is obtained by using a preset center point detection strategy and a preset size regression strategy based on the top view, and center point coordinate information and state information of each target object in the target detection area are obtained based on the target thermodynamic diagram.
Specifically, after the low-dimensional feature representation (i.e., the feature tensor of the preset dimension) is obtained, the feature tensor of the preset dimension can be generated into a top view using a correlation algorithm or a correlation neural network model. After obtaining the top view, the top view can be further processed by using a preset center point detection strategy and a preset size regression strategy, so as to obtain a target thermodynamic diagram. The target heat map refers to a two-dimensional image representing the existence probability of a target by color or brightness, and each local maximum value in the target heat map corresponds to the center point of one target object (i.e. a vehicle).
How the target thermodynamic diagram is obtained using a preset center point detection strategy and a preset size regression strategy based on the top view is described in detail below.
As a possible implementation manner, in some embodiments, the target thermodynamic diagram is obtained by utilizing a preset center point detection strategy and a preset size regression strategy based on a top view, and the method comprises the steps of identifying at least one preset key point in the top view by utilizing the preset center point strategy, predicting three-dimensional boundary box information of a target object corresponding to each preset key point by utilizing the preset size regression strategy based on the at least one preset key point, and obtaining the target thermodynamic diagram by utilizing a preset dynamic Gaussian function radius strategy based on the three-dimensional boundary box information.
The preset key point is the center point of the target object. The three-dimensional bounding box is a rectangular box in three-dimensional space that accurately describes the position and size of the target object.
Specifically, the image-based key point detection technology (i.e., a preset key point detection strategy) is used to effectively identify the center point (i.e., at least one preset key point) of each object positioned in the top view, then, for each identified object center point, a size regression operation is further performed based on a preset size regression strategy to calculate the complete three-dimensional bounding box information of the target object, including the coordinate information (i.e., the center position) and the state information (i.e., the length, the width, the height, the heading angle, the speed, etc.) of the center point of the target object, and finally, to solve the problem that the occupied area of the target object such as a vehicle in the map view is small, the forward supervision of the thermodynamic diagram can be increased by using a preset dynamic gaussian function radius strategy according to the three-dimensional bounding box information, so as to generate the target thermodynamic diagram.
It will be appreciated that the radius of the gaussian function may be dynamically adjusted according to the size and shape of the target object in order to better accommodate different sized target objects. For example, for each predicted target object, a gaussian distributed hot spot region is generated on the thermodynamic diagram according to three-dimensional bounding box information, and the radius of the hot spot region is dynamically adjusted according to the size of the target object.
Further, in some embodiments, when the coordinate information and the state information of the central point of each target object in the target detection area are obtained based on the target thermodynamic diagram, the method further comprises the steps of carrying out feature extraction operation on the three-dimensional central points of each face of each three-dimensional boundary frame based on the three-dimensional boundary frame information of each target object in the target thermodynamic diagram to obtain feature values of each three-dimensional central point, carrying out stacking operation on the feature values of each three-dimensional central point based on a bilinear interpolation strategy to obtain feature vectors, and inputting the feature vectors into a preset neural network to obtain optimized three-dimensional boundary frame information of each target object.
Specifically, by performing a feature extraction operation on the three-dimensional center point of each face of each three-dimensional bounding box, a feature value of each three-dimensional center point can be obtained. Then, stacking the characteristic values of each three-dimensional center point through a bilinear interpolation strategy to form a characteristic vector, then inputting the characteristic vector into a preset neural network (such as a multi-layer perceptron), and obtaining optimized three-dimensional boundary frame information of each target object through network processing and optimization.
For further understanding of the object detection method based on color point cloud according to the embodiments of the present application, the following description is further provided with reference to fig. 2.
As shown in fig. 2, the color point cloud-based target detection method may include the following steps:
Step S201, laser radar data and RGB image data are input.
Step S202, the laser radar data are projected onto the RGB image data through the internal parameters and the external parameters of the camera and the laser radar, and RGB values of RGB image pixels corresponding to the laser radar data are assigned to the laser radar data to generate color point cloud data.
In step S203, the color point cloud data is subjected to voxel processing, that is, the three-dimensional space is divided into small voxel units, each voxel unit contains a plurality of point cloud data points, and structured data is obtained. For example, in an autopilot scenario, the space around the vehicle may be divided into voxels of a size of 0.1 meter or 0.2 meter or the like in side length.
And S204, carrying out local feature extraction on the point cloud data in each voxel by utilizing a three-dimensional backbone network so as to acquire the feature representation of each voxel. Then, further processing and aggregation are carried out on the voxel characteristics through a series of three-dimensional convolution layers, the receptive field is gradually enlarged, and higher-level semantic characteristics are extracted, so that a three-dimensional characteristic diagram containing rich semantic information is obtained.
In step S205, the three-dimensional feature map is flattened into a plan view, and a center point of the vehicle is found using an image-based key point detector, and a thermodynamic diagram of the vehicle type is generated by a thermodynamic diagram regression method. For each detected vehicle center point, other attributes of the vehicle (namely three-dimensional boundary box information) including three-dimensional dimensions (length, width and height), three-dimensional directions (course angle) and speeds are regressed from the point characteristics of the center position. For example, the information that the length of the vehicle is 4.5 meters, the width is 1.8 meters, the height is 1.5 meters, the current heading angle is 30 degrees, and the speed is 10 meters/second is predicted by the regression head.
Step S206, extracting a point feature from the three-dimensional center of each face of the predicted frame according to the predicted vehicle boundary frame information, extracting a feature from the feature map output by the backbone network by using bilinear interpolation for each point, inputting the point features into the fully-connected network for refinement, obtaining more accurate vehicle position and size information, and simultaneously predicting a confidence score, wherein the confidence score represents 3D IoU (Intersection over Union,3D cross-point ratio) between the predicted result and the true value. 3D IoU is an index for evaluating the degree of overlap between a prediction bounding box and a real bounding box in three-dimensional space.
In step S207, information such as the center point position, the three-dimensional size, the three-dimensional direction, the speed, and the confidence score of each detected vehicle is output. Such information may be used by the autopilot system for further decisions and control, such as planning the travel path of the vehicle, avoiding collisions, etc.
According to the target detection method based on the color point cloud, the color point cloud data can be obtained through laser radar data and RGB image data based on the target detection area, the color point cloud data are subjected to voxelization to obtain structural data, a preset three-dimensional encoder network is utilized to compress and extract features of the structural data to obtain feature tensors of preset dimensions, a top view is generated based on the feature tensors of the preset dimensions, a target thermodynamic diagram is obtained based on the top view by utilizing a preset center point detection strategy and a preset size regression strategy, and then center point coordinate information and state information of each target object in the target detection area are obtained. Therefore, the problem of low target detection precision under the complex environment only depending on the laser radar data is solved by fusing the camera data and the laser radar data to perform target detection, and the accuracy and the robustness of target detection are improved.
Next, a color point cloud-based object detection device according to an embodiment of the present application is described with reference to the accompanying drawings.
Fig. 3 is a block schematic diagram of a color point cloud-based object detection apparatus according to an embodiment of the present application.
As shown in fig. 3, the color point cloud-based object detection apparatus 10 includes a first obtaining module 100, a processing module 200, and a second obtaining module 300.
The first obtaining module 100 is configured to obtain laser radar data and RGB image data of a target detection area, and obtain color point cloud data based on the laser radar data and the RGB image data;
The processing module 200 is configured to voxel process the color point cloud data to obtain structured data, and perform compression and feature extraction operations on the structured data by using a preset three-dimensional encoder network to obtain a feature tensor of a preset dimension;
the second obtaining module 300 is configured to generate a top view based on the feature tensor of the preset dimension, obtain a target thermodynamic diagram based on the top view by using a preset center point detection strategy and a preset size regression strategy, and obtain center point coordinate information and state information of each target object in the target detection area based on the target thermodynamic diagram.
Optionally, in some embodiments, the first obtaining module 100 is specifically configured to:
Acquiring a first internal parameter and a first external parameter of the laser radar, and acquiring a second internal parameter and a second external parameter of the camera;
Based on the first internal parameter, the first external parameter, the second internal parameter and the second external parameter, the three-dimensional coordinates corresponding to the laser radar data are projected to a two-dimensional coordinate system corresponding to the RGB image data, and based on a preset pixel-level RGB assignment strategy, RGB values of corresponding pixels are given to each laser radar data projected to the two-dimensional coordinate system corresponding to the RGB image data, so that color point cloud data are generated.
Optionally, in some embodiments, the processing module 200 is specifically configured to:
Based on a preset spatial resolution condition, discretizing a three-dimensional coordinate system of color point cloud data into a plurality of cube units arranged according to a preset arrangement rule, and obtaining structured data based on the plurality of cube units arranged according to the preset arrangement rule.
Optionally, in some embodiments, the second obtaining module 300 is specifically configured to:
identifying at least one preset key point in the top view by using a preset center point strategy;
predicting three-dimensional boundary frame information of a target object corresponding to each preset key point by utilizing a preset size regression strategy based on at least one preset key point;
And obtaining a target thermodynamic diagram by utilizing a preset dynamic Gaussian function radius strategy based on the three-dimensional bounding box information.
Optionally, in some embodiments, when obtaining the center point coordinate information and the state information of each target object in the target detection area based on the target thermodynamic diagram, the second obtaining module 300 is further configured to:
based on the three-dimensional boundary frame information of each target object in the target thermodynamic diagram, carrying out feature extraction operation on the three-dimensional center point of each surface of each three-dimensional boundary frame to obtain a feature value of each three-dimensional center point;
Based on a bilinear interpolation strategy, stacking the characteristic values of each three-dimensional center point to obtain characteristic vectors, and inputting the characteristic vectors into a preset neural network to obtain optimized three-dimensional boundary box information of each target object.
It should be noted that the foregoing explanation of the embodiment of the color point cloud-based target detection method is also applicable to the color point cloud-based target detection device of the embodiment, and will not be repeated herein.
According to the target detection device based on the color point cloud, the color point cloud data can be obtained through laser radar data and RGB image data based on the target detection area, the color point cloud data are subjected to voxelization to obtain structural data, a preset three-dimensional encoder network is utilized to compress and extract features of the structural data to obtain feature tensors of preset dimensions, a top view is generated based on the feature tensors of the preset dimensions, a target thermodynamic diagram is obtained based on the top view by utilizing a preset center point detection strategy and a preset size regression strategy, and then center point coordinate information and state information of each target object in the target detection area are obtained. Therefore, the problem of low target detection precision under the complex environment only depending on the laser radar data is solved by fusing the camera data and the laser radar data to perform target detection, and the accuracy and the robustness of target detection are improved.
Fig. 4 is a schematic structural diagram of an electronic device according to an embodiment of the present application. The electronic device may include:
memory 401, processor 402, and a computer program stored on memory 401 and executable on processor 402.
The processor 402 implements the color point cloud-based target detection method provided in the above embodiment when executing a program.
Further, the electronic device further includes:
A communication interface 403 for communication between the memory 401 and the processor 402.
A memory 401 for storing a computer program executable on the processor 402.
Memory 401 may include high-speed RAM (Random Access Memory ) memory, and may also include non-volatile memory, such as at least one disk memory.
If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 may be connected to each other by a bus and perform communication with each other. The bus may be an ISA (Industry Standard Architecture ) bus, a PCI (PERIPHERAL COMPONENT INTERCONNECT, external device interconnect) bus, or EISA (Extended Industry Standard Architecture ) bus, among others. The buses may be divided into address buses, data buses, control buses, etc. For ease of illustration, only one thick line is shown in fig. 4, but not only one bus or one type of bus.
Alternatively, in a specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a chip, the memory 401, the processor 402, and the communication interface 403 may perform communication with each other through internal interfaces.
The processor 402 may be a CPU (Central Processing Unit ) or an ASIC (Application SPECIFIC INTEGRATED Circuit, application specific integrated Circuit) or one or more integrated circuits configured to implement embodiments of the present application.
The embodiment of the application also provides a computer readable storage medium, on which a computer program is stored, which when being executed by a processor, implements the target detection method based on color point cloud as above.
Furthermore, the terms "first," "second," and the like, are used for descriptive purposes only and are not to be construed as indicating or implying a relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defining "a first" or "a second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "plurality" means at least two, for example, two, three, etc., unless specifically defined otherwise.
In the description of the present specification, a description referring to terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples," etc., means that a particular feature, structure, material, or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms are not necessarily directed to the same embodiment or example. Furthermore, the particular features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, the different embodiments or examples described in this specification and the features of the different embodiments or examples may be combined and combined by those skilled in the art without contradiction.
While embodiments of the present application have been shown and described above, it will be understood that the above embodiments are illustrative and not to be construed as limiting the application, and that variations, modifications, alternatives and variations may be made to the above embodiments by one of ordinary skill in the art within the scope of the application.

Claims (10)

1. The target detection method based on the color point cloud is characterized by comprising the following steps of:
acquiring laser radar data and RGB image data of a target detection area, and acquiring color point cloud data based on the laser radar data and the RGB image data;
Carrying out voxelization on the color point cloud data to obtain structured data, and carrying out compression and feature extraction operation on the structured data by utilizing a preset three-dimensional encoder network to obtain a feature tensor with preset dimension;
Generating a top view based on the characteristic tensor of the preset dimension, obtaining a target thermodynamic diagram by utilizing a preset center point detection strategy and a preset size regression strategy based on the top view, and obtaining center point coordinate information and state information of each target object in the target detection area based on the target thermodynamic diagram.
2. The method of claim 1, wherein the deriving color point cloud data based on the lidar data and RGB image data comprises:
Acquiring a first internal parameter and a first external parameter of the laser radar, and acquiring a second internal parameter and a second external parameter of the camera;
Based on the first internal parameter, the first external parameter, the second internal parameter and the second external parameter, projecting the three-dimensional coordinate corresponding to the laser radar data to the two-dimensional coordinate system corresponding to the RGB image data, and based on a preset pixel-level RGB assignment strategy, assigning RGB values of corresponding pixels to each laser radar data projected to the two-dimensional coordinate system corresponding to the RGB image data, and generating the color point cloud data.
3. The method of claim 1, wherein said voxelizing the color point cloud data to obtain structured data, comprising:
and dispersing the three-dimensional coordinate system of the color point cloud data into a plurality of cube units which are arranged according to a preset arrangement rule based on a preset spatial resolution condition, and obtaining the structured data based on the plurality of cube units which are arranged according to the preset arrangement rule.
4. The method of claim 1, wherein the obtaining the target thermodynamic diagram based on the top view using a preset center point detection strategy and a preset size regression strategy comprises:
identifying at least one preset key point in the top view by utilizing the preset center point strategy;
Predicting three-dimensional boundary box information of a target object corresponding to each preset key point by utilizing the preset size regression strategy based on the at least one preset key point;
And obtaining the target thermodynamic diagram by utilizing a preset dynamic Gaussian function radius strategy based on the three-dimensional boundary box information.
5. The method of claim 4, wherein when deriving the center point coordinate information and the status information for each target object in the target detection area based on the target thermodynamic diagram, further comprising:
based on the three-dimensional boundary frame information of each target object in the target thermodynamic diagram, carrying out feature extraction operation on the three-dimensional center point of each surface of each three-dimensional boundary frame to obtain a feature value of each three-dimensional center point;
and stacking the characteristic values of each three-dimensional center point based on a bilinear interpolation strategy to obtain characteristic vectors, and inputting the characteristic vectors into a preset neural network to obtain optimized three-dimensional boundary box information of each target object.
6. Target detection device based on color point cloud, characterized by comprising:
The first acquisition module is used for acquiring laser radar data and RGB image data of the target detection area and acquiring color point cloud data based on the laser radar data and the RGB image data;
The processing module is used for carrying out voxelization on the color point cloud data to obtain structured data, and carrying out compression and feature extraction operation on the structured data by utilizing a preset three-dimensional encoder network to obtain a feature tensor with preset dimension;
the second obtaining module is configured to generate a top view based on the feature tensor of the preset dimension, obtain a target thermodynamic diagram based on the top view by using a preset center point detection strategy and a preset size regression strategy, and obtain center point coordinate information and state information of each target object in the target detection area based on the target thermodynamic diagram.
7. The apparatus of claim 6, wherein the first obtaining module is specifically configured to:
Acquiring a first internal parameter and a first external parameter of the laser radar, and acquiring a second internal parameter and a second external parameter of the camera;
Based on the first internal parameter, the first external parameter, the second internal parameter and the second external parameter, projecting the three-dimensional coordinate corresponding to the laser radar data to the two-dimensional coordinate system corresponding to the RGB image data, and based on a preset pixel-level RGB assignment strategy, assigning RGB values of corresponding pixels to each laser radar data projected to the two-dimensional coordinate system corresponding to the RGB image data, and generating the color point cloud data.
8. The apparatus of claim 6, wherein the processing module is specifically configured to:
and dispersing the three-dimensional coordinate system of the color point cloud data into a plurality of cube units which are arranged according to a preset arrangement rule based on a preset spatial resolution condition, and obtaining the structured data based on the plurality of cube units which are arranged according to the preset arrangement rule.
9. An electronic device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, the processor executing the program to implement the color point cloud-based object detection method according to any one of claims 1-5.
10. A computer readable storage medium having stored thereon a computer program, characterized in that the program is executed by a processor for implementing the color point cloud based object detection method according to any of claims 1-5.
CN202510819580.5A 2025-06-18 2025-06-18 Target detection method and device based on color point cloud, electronic equipment and medium Pending CN120807875A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
CN202510819580.5A CN120807875A (en) 2025-06-18 2025-06-18 Target detection method and device based on color point cloud, electronic equipment and medium

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
CN202510819580.5A CN120807875A (en) 2025-06-18 2025-06-18 Target detection method and device based on color point cloud, electronic equipment and medium

Publications (1)

Publication Number Publication Date
CN120807875A true CN120807875A (en) 2025-10-17

Family

ID=97320295

Family Applications (1)

Application Number Title Priority Date Filing Date
CN202510819580.5A Pending CN120807875A (en) 2025-06-18 2025-06-18 Target detection method and device based on color point cloud, electronic equipment and medium

Country Status (1)

Country Link
CN (1) CN120807875A (en)

Similar Documents

Publication Publication Date Title
CN113819890B (en) Distance measuring method, distance measuring device, electronic equipment and storage medium
CN113284163B (en) Three-dimensional target self-adaptive detection method and system based on vehicle-mounted laser radar point cloud
US12315165B2 (en) Object detection method, object detection device, terminal device, and medium
KR102029850B1 (en) Object detecting apparatus using camera and lidar sensor and method thereof
CN113408324A (en) Target detection method, device and system and advanced driving assistance system
CN113658257B (en) Unmanned equipment positioning method, device, equipment and storage medium
CN115272416A (en) Vehicle and pedestrian detection tracking method and system based on multi-source sensor fusion
CN113761999A (en) Target detection method and device, electronic equipment and storage medium
US12293593B2 (en) Object detection method, object detection device, terminal device, and medium
CN112613378A (en) 3D target detection method, system, medium and terminal
CN115147333A (en) Target detection method and device
CN115035492B (en) Vehicle identification method, device, equipment and storage medium
CN114119992A (en) Multi-mode three-dimensional target detection method and device based on image and point cloud fusion
CN114913519A (en) A 3D target detection method, device, electronic device and storage medium
CN117727026A (en) A multi-modal fusion target detection method based on unified BEV representation
Liu et al. Vehicle-related distance estimation using customized YOLOv7
CN118096834B (en) YOLO-based multi-sensor fusion dynamic object tracking method
CN114972492A (en) A bird's-eye view-based pose determination method, device and computer storage medium
CN120411714A (en) AI visual target detection method based on multimodal feature fusion
CN116311114A (en) A drivable area generation method, device, electronic equipment and storage medium
Venugopala Comparative study of 3D object detection frameworks based on LiDAR data and sensor fusion techniques
CN116778262B (en) Three-dimensional target detection method and system based on virtual point cloud
WO2024239605A1 (en) Attack detection method and apparatus for autonomous driving system, device, and storage medium
CN114140659B (en) A social distance monitoring method based on human body detection from the perspective of drones
CN112464905B (en) 3D target detection method and device

Legal Events

Date Code Title Description
PB01 Publication
PB01 Publication
SE01 Entry into force of request for substantive examination
SE01 Entry into force of request for substantive examination