CN116543369A - Method and device for identifying drivable area, electronic equipment and storage medium - Google Patents
Method and device for identifying drivable area, electronic equipment and storage medium Download PDFInfo
- Publication number
- CN116543369A CN116543369A CN202310504926.3A CN202310504926A CN116543369A CN 116543369 A CN116543369 A CN 116543369A CN 202310504926 A CN202310504926 A CN 202310504926A CN 116543369 A CN116543369 A CN 116543369A
- Authority
- CN
- China
- Prior art keywords
- feature map
- image
- target
- attention
- sample
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
- G06V20/58—Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/52—Scale-space analysis, e.g. wavelet analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
- G06V10/806—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
- G06V20/58—Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads
- G06V20/582—Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads of traffic signs
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
- G06V20/588—Recognition of the road, e.g. of lane markings; Recognition of the vehicle driving pattern in relation to the road
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02T—CLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO TRANSPORTATION
- Y02T10/00—Road transport of goods or passengers
- Y02T10/10—Internal combustion engine [ICE] based vehicles
- Y02T10/40—Engine management systems
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Evolutionary Computation (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Artificial Intelligence (AREA)
- Computing Systems (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Software Systems (AREA)
- Databases & Information Systems (AREA)
- Medical Informatics (AREA)
- Biomedical Technology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Molecular Biology (AREA)
- General Engineering & Computer Science (AREA)
- Mathematical Physics (AREA)
- Image Analysis (AREA)
Abstract
Description
技术领域technical field
本申请涉及智能驾驶技术领域,尤其涉及一种可行驶区域识别方法、装置、电子设备及存储介质。The present application relates to the technical field of intelligent driving, and in particular to a drivable area identification method, device, electronic equipment and storage medium.
背景技术Background technique
目前,检测车辆行驶过程中的场景信息是辅助驾驶系统和自动驾驶系统中的关键,场景信息的检测包括车道线、交通标识、道路标识及可行驶区域等,其中,可行驶区域检测在辅助驾驶系统和自动驾驶系统中发挥着重要作用,在车辆行驶过程中,为车辆提供可行驶路线,避开不可行驶区域,不可行驶区域主要包括道路上的静态障碍物,例如,锥桶、护栏等所在的区域,以及动态障碍物,例如,车辆、行人等所在的区域。At present, the detection of scene information during vehicle driving is the key to assisted driving systems and automatic driving systems. The detection of scene information includes lane lines, traffic signs, road signs, and drivable areas. The system and the automatic driving system play an important role. During the driving process of the vehicle, it provides the vehicle with a drivable route and avoids the undrivable area. The undrivable area mainly includes static obstacles on the road, such as cones, guardrails, etc. area, as well as the area where dynamic obstacles, such as vehicles and pedestrians, are located.
相关技术中,在识别可行驶区域时,通常对采集的图片,使用复杂的编码器提取特征,获得四倍下采样特征图。例如,高分辨率网络(High-Resolution Net,HRNet),再将四倍下采样特征图进行放大处理,使得输出图片恢复到输入图片的尺寸,随后获得可行驶区域。In related technologies, when identifying a drivable area, a complex encoder is usually used to extract features from the collected pictures to obtain a quadruple downsampled feature map. For example, the high-resolution network (High-Resolution Net, HRNet), and then enlarge the four times downsampled feature map, so that the output image can be restored to the size of the input image, and then the drivable area can be obtained.
然而,采用上述方法时,由于复杂的编码器的内存访问代价较大,因此图形处理器(graphics processing unit,GPU)的计算效率较低,限制了识别速度。However, when the above method is adopted, the computational efficiency of the graphics processing unit (GPU) is low due to the high memory access cost of the complex encoder, which limits the recognition speed.
另外,为了提升识别速度,直接采用放大处理恢复图片的尺寸,由于四倍下采样特征图的像素点较小,对象不能够观察清楚像素点,因此若直接对四倍下采样特征图进行放大处理,会造成识别不准确的情况。In addition, in order to improve the recognition speed, the size of the image is directly restored by zooming in. Since the pixels of the quadruple downsampling feature map are small, the object cannot observe the pixels clearly. Therefore, if the quadruple downsampling feature map is directly enlarged , resulting in inaccurate recognition.
例如,当四倍下采样特征图存在几个可行驶区域的像素点未被识别成可行驶区域的情况时,进行放大处理后,像素点会变大,对象能够观察到有一小块区域没有被识别成可行驶区域,导致出现空洞问题。For example, when the quadruple downsampling feature map has several pixels in the drivable area that are not recognized as drivable areas, after the zoom-in process, the pixels will become larger, and the object can observe that there is a small area that is not recognized as a drivable area. Identified as a drivable area, resulting in a hole problem.
因此,相关技术中的可行驶区域识别的结果准确率和识别效率都有待提高。Therefore, the result accuracy and recognition efficiency of drivable area recognition in the related art need to be improved.
发明内容Contents of the invention
本申请实施例提供一种可行驶区域识别方法、装置、电子设备及存储介质,以提高可行驶区域识别的结果准确率和识别效率。Embodiments of the present application provide a drivable area identification method, device, electronic device, and storage medium, so as to improve the result accuracy and identification efficiency of drivable area identification.
本申请实施例提供的具体技术方案如下:The specific technical scheme that the embodiment of the present application provides is as follows:
第一方面,提供一种可行驶区域识别方法,包括:In the first aspect, a method for identifying a drivable area is provided, including:
响应于驾驶对象触发的可行驶区域识别请求,获取当前待识别的目标道路图像,目标道路图像具有目标图像尺度;Responding to a drivable area identification request triggered by a driving object, acquiring a current target road image to be identified, where the target road image has a target image scale;
按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图;According to the preset scales of each candidate image, feature extraction is performed on the target road image to obtain the corresponding road feature map;
针对各道路特征图,分别执行以下操作:基于一个道路特征图包含的各像素点各自的图像位置信息,结合一个道路特征图包含的图像通道信息,对道路特征图进行注意力加权处理,获得相应的注意力特征图;For each road feature map, perform the following operations: Based on the image position information of each pixel contained in a road feature map, combined with the image channel information contained in a road feature map, perform attention weighting processing on the road feature map to obtain the corresponding The attention feature map of
基于目标图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,并基于至少一个目标特征图,获得目标车辆的可行驶区域。Based on the target image scale, image scale adjustment is performed on each of the obtained attention feature maps to obtain at least one target feature map, and based on the at least one target feature map, a drivable area of the target vehicle is obtained.
第二方面,提供一种可行驶区域识别装置,包括:In a second aspect, a drivable area identification device is provided, including:
获取模块,用于响应于驾驶对象触发的可行驶区域识别请求,获取当前待识别的目标道路图像,目标道路图像具有目标图像尺度;An acquisition module, configured to acquire a current target road image to be identified in response to a drivable area identification request triggered by a driving object, where the target road image has a target image scale;
提取模块,用于按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图;The extraction module is used to extract the features of the target road image according to the preset scales of each candidate image, and obtain the corresponding road feature map;
第一处理模块,用于针对各道路特征图,分别执行以下操作:基于一个道路特征图包含的各像素点各自的图像位置信息,结合一个道路特征图包含的图像通道信息,对道路特征图进行注意力加权处理,获得相应的注意力特征图;The first processing module is configured to perform the following operations on each road feature map: based on the image position information of each pixel contained in a road feature map, combined with the image channel information contained in a road feature map, the road feature map is processed Attention weighting processing to obtain the corresponding attention feature map;
第二处理模块,用于基于目标图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,并基于至少一个目标特征图,获得目标车辆的可行驶区域。The second processing module is configured to perform image scale adjustment on each of the obtained attention feature maps based on the target image scale, obtain at least one target feature map, and obtain the drivable area of the target vehicle based on the at least one target feature map.
可选的,按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图时,提取模块用于:Optionally, according to the preset scales of each candidate image, feature extraction is performed on the target road image respectively, and when the corresponding road feature map is obtained, the extraction module is used for:
针对预设的各候选图像尺度,分别执行以下操作:For each preset candidate image scale, perform the following operations:
基于一个候选图像尺度,对目标道路图像进行第一特征提取,获得第一特征图,第一特征图具有第一图像通道数;Based on a candidate image scale, performing first feature extraction on the target road image to obtain a first feature map, the first feature map has a first number of image channels;
按照预设的各候选图像感受野,分别对第一特征图进行第二特征提取,获得相应的第二特征图,并对获得的各第二特征图进行叠加,获得叠加特征图;其中,每个候选图像感受野表征:相应第二特征图包含的各像素点映射至目标道路图像上的区域大小;According to the preset receptive field of each candidate image, the second feature extraction is performed on the first feature map to obtain the corresponding second feature map, and the obtained second feature maps are superimposed to obtain the superimposed feature map; wherein, each A representation of the receptive field of a candidate image: each pixel contained in the corresponding second feature map is mapped to the area size on the target road image;
基于第一图像通道数,对叠加特征图进行图像通道调整,获得一个候选图像尺度对应的道路特征图。Based on the number of first image channels, image channel adjustment is performed on the superposition feature map to obtain a road feature map corresponding to a candidate image scale.
可选的,基于一个道路特征图包含的各像素点各自的图像位置信息,结合一个道路特征图包含的图像通道信息,对道路特征图进行注意力加权处理,获得相应的注意力特征图时,第一处理模块用于:Optionally, based on the respective image position information of each pixel contained in a road feature map, combined with the image channel information contained in a road feature map, the road feature map is subjected to attention weighting processing, and when the corresponding attention feature map is obtained, The first processing module is used for:
基于一个道路特征图包含的各像素点各自的图像位置信息,确定一个道路特征图包含的各像素点各自的位置注意力权重,其中,位置注意力权重表征在识别可行驶区域时,相应像素点的图像位置的重要程度;Based on the respective image position information of each pixel contained in a road feature map, determine the respective position attention weights of each pixel contained in a road feature map, where the position attention weight represents the corresponding pixel points when identifying the drivable area The importance of the image position;
基于一个道路特征图包含的图像通道信息,确定一个道路特征图包含的各图像通道各自的通道注意力权重,其中,通道注意力权重表征在识别可行驶区域时,相应图像通道的重要程度;Based on the image channel information contained in a road feature map, determine the respective channel attention weights of each image channel contained in a road feature map, wherein the channel attention weight represents the importance of the corresponding image channel when identifying the drivable area;
基于获得的各位置注意力权重和通道注意力权重,对一个道路特征图进行加权,获得一个道路特征图对应的注意力特征图。Based on the obtained attention weights of each position and channel attention weights, a road feature map is weighted to obtain an attention feature map corresponding to a road feature map.
可选的,基于目标图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,第二处理模块用于:Optionally, based on the target image scale, image scale adjustment is performed on each of the obtained attention feature maps to obtain at least one target feature map, and the second processing module is used for:
基于获得的各注意力特征图各自的图像尺度大小,对各注意力特征图进行融合更新,获得至少一个融合特征图;Based on the respective image scales of the obtained attention feature maps, fusion and update are performed on each attention feature map to obtain at least one fusion feature map;
基于预设的目标图像通道数,分别对至少一个融合特征图进行图像通道调整,获得相应的中间特征图;Based on the preset number of target image channels, image channel adjustment is performed on at least one fusion feature map to obtain a corresponding intermediate feature map;
基于目标图像尺度,对获得的至少一个中间特征图进行图像尺度调整,获得相应的目标特征图。Based on the target image scale, image scale adjustment is performed on at least one obtained intermediate feature map to obtain a corresponding target feature map.
可选的,基于获得的各注意力特征图各自的图像尺度大小,对各注意力特征图进行融合更新,获得至少一个融合特征图时,第二处理模块还用于:Optionally, based on the respective image scales of the obtained attention feature maps, the attention feature maps are fused and updated, and when at least one fusion feature map is obtained, the second processing module is also used for:
按照获得的各注意力特征图各自的图像尺度大小,将各注意力特征图进行排序,获得排序结果;According to the respective image scales of the obtained attention feature maps, sort the attention feature maps to obtain the sorting results;
基于排序结果,依次读取每两个相邻的注意力特征图进行融合,直到读取完毕为止,其中,一次融合过程包括:Based on the sorting results, each two adjacent attention feature maps are read sequentially for fusion until the reading is completed. A fusion process includes:
读取相邻的两个注意力特征图,并对两个注意力特征图进行融合,获得融合特征图;Read two adjacent attention feature maps, and fuse the two attention feature maps to obtain a fusion feature map;
保存融合特征图,以及将融合特征图作为新的注意力特征图,代替两个注意力特征图,加入排序结果中。Save the fusion feature map, and use the fusion feature map as a new attention feature map to replace the two attention feature maps and add it to the sorting result.
可选的,基于至少一个目标特征图,获得目标车辆的可行驶区域时,第二处理模块还用于:Optionally, when obtaining the drivable area of the target vehicle based on at least one target feature map, the second processing module is also used for:
若至少一个目标特征图包括一个目标特征图,则将一个目标特征图作为结果特征图;If at least one target feature map includes a target feature map, then using a target feature map as a result feature map;
若至少一个目标特征图包括多个目标特征图,则对多个目标特征图进行图像融合,获得结果特征图;If at least one target feature map includes multiple target feature maps, performing image fusion on the multiple target feature maps to obtain a resultant feature map;
将结果特征图中包含的各像素点的值与预设的可行驶区域标签值进行比较,确定结果特征图中属于可行驶区域的各目标像素点;Comparing the value of each pixel contained in the result feature map with the preset drivable area label value, and determining each target pixel point belonging to the drivable area in the result feature map;
将目标道路图像中,与各目标像素点各自的图像位置对应的各像素点进行标示,获得包含可行驶区域的图像。In the target road image, each pixel point corresponding to the respective image position of each target pixel point is marked to obtain an image including a drivable area.
可选的,可行驶区域是通过将目标道路图像,输入目标识别模型获得的,所述装置还包括训练模块,训练模块用于:Optionally, the drivable area is obtained by inputting the target road image into the target recognition model, and the device also includes a training module, which is used for:
基于训练样本集对待训练的识别模型进行迭代训练,获得目标识别模型;每个训练样本包括:样本道路的图像数据,其中,每个迭代过程执行以下操作:Perform iterative training on the recognition model to be trained based on the training sample set to obtain the target recognition model; each training sample includes: image data of the sample road, wherein each iteration process performs the following operations:
按照预设的各样本图像尺度,分别对选取的训练样本进行特征提取,获得相应的样本特征图,训练样本具有样本图像尺度;According to the preset image scales of each sample, feature extraction is performed on the selected training samples to obtain corresponding sample feature maps, and the training samples have sample image scales;
针对各样本特征图,分别执行以下操作:基于一个样本特征图包含的各像素点各自的样本图像位置信息,结合一个样本特征图包含的样本图像通道信息,对样本特征图进行注意力加权处理,获得相应的样本注意力特征图;For each sample feature map, perform the following operations respectively: Based on the respective sample image position information of each pixel contained in a sample feature map, combined with the sample image channel information contained in a sample feature map, perform attention weighting processing on the sample feature map, Obtain the corresponding sample attention feature map;
基于样本图像尺度,分别对获得的各样本注意力特征图进行图像尺度调整,获得至少一个目标样本特征图;Based on the sample image scale, image scale adjustment is performed on the obtained attention feature maps of each sample to obtain at least one target sample feature map;
基于至少一个目标样本特征图,获得训练样本的可行驶区域预测结果,并基于可行驶区域预测结果对应的损失值进行调参。Based on at least one target sample feature map, the drivable area prediction result of the training sample is obtained, and the parameters are adjusted based on the loss value corresponding to the drivable area prediction result.
第三方面,提供一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现上述第一方面任一项所述方法的步骤。In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and operable on the processor, and the processor implements any one of the above-mentioned first aspects when executing the program. method steps.
第四方面,提供一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现上述第一方面任一项所述方法的步骤。In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the methods described in the above-mentioned first aspect are implemented.
本申请实施例中,在驾驶对象触发可行驶区域识别请求时,服务器响应于驾驶对象触发的可行驶区域识别请求,获取当前待识别的目标道路图像,再按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图,然后,对各道路特征图在空间位置维度和通道维度分别进行注意力加权处理,获得各注意力特征图,最后,基于目标道路图像的图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,并基于至少一个目标特征图,获得目标车辆的可行驶区域。In the embodiment of the present application, when the driving object triggers a drivable area identification request, the server responds to the drivable area identification request triggered by the driving object, obtains the current target road image to be identified, and then according to the preset candidate image scales, respectively Feature extraction is performed on the target road image to obtain the corresponding road feature map. Then, the attention weighting process is performed on each road feature map in the spatial position dimension and the channel dimension to obtain each attention feature map. Finally, based on the target road image The image scale is to adjust the image scale of each obtained attention feature map to obtain at least one target feature map, and obtain the drivable area of the target vehicle based on the at least one target feature map.
这样,获得多尺度的各道路特征图,对各道路特征图在空间位置维度和通道维度分别进行加权处理,并对加权后的各道路特征图进行融合操作,能够更好地平衡多尺度道路特征图的空间信息和语义信息,避免了空洞问题,提高了可行驶区域识别的结果准确率和识别效率,并且目标特征图和目标道路图像的图像尺度保持一致,能够更准确的获得目标车辆的可行驶区域。In this way, obtaining multi-scale road feature maps, weighting each road feature map in the spatial position dimension and channel dimension, and performing fusion operations on the weighted road feature maps can better balance multi-scale road features. The spatial information and semantic information of the map avoid the hole problem, improve the accuracy and recognition efficiency of the drivable area recognition result, and the image scale of the target feature map and the target road image are consistent, which can obtain the drivable area of the target vehicle more accurately. driving area.
附图说明Description of drawings
图1为本申请实施例中可能的应用场景示意图;FIG. 1 is a schematic diagram of a possible application scenario in the embodiment of the present application;
图2为本申请实施例中目标识别模型一轮训练的流程示意图;FIG. 2 is a schematic flow diagram of one round of training of the target recognition model in the embodiment of the present application;
图3为本申请实施例中编码器的示意图;FIG. 3 is a schematic diagram of an encoder in an embodiment of the present application;
图4为本申请实施例中高性能GPU模块层的示意图;Fig. 4 is the schematic diagram of high-performance GPU module layer in the embodiment of the present application;
图5为本申请实施例中空间注意力模块和通道注意力模块的示意图;5 is a schematic diagram of a spatial attention module and a channel attention module in an embodiment of the present application;
图6为本申请实施例中可行驶区域识别方法的流程示意图;FIG. 6 is a schematic flow chart of a drivable area identification method in an embodiment of the present application;
图7为本申请实施例中获得道路特征图的流程示意图;FIG. 7 is a schematic flow diagram of obtaining a road feature map in an embodiment of the present application;
图8为本申请实施例中获得叠加特征图的示意图;FIG. 8 is a schematic diagram of obtaining a superposition feature map in an embodiment of the present application;
图9为本申请实施例中获得一个道路特征图对应的注意力特征图的流程示意图;FIG. 9 is a schematic flow diagram of obtaining an attention feature map corresponding to a road feature map in an embodiment of the present application;
图10为本申请实施例中获得注意力特征图的示意图;FIG. 10 is a schematic diagram of obtaining an attention feature map in an embodiment of the present application;
图11为本申请实施例获得至少一个目标特征图的流程示意图;FIG. 11 is a schematic flow diagram of obtaining at least one target feature map according to an embodiment of the present application;
图12为本申请实施例获得至少一个融合特征图的流程示意图;FIG. 12 is a schematic flow diagram of obtaining at least one fusion feature map according to an embodiment of the present application;
图13为本申请实施例中获得至少一个融合特征图的示意图;Fig. 13 is a schematic diagram of obtaining at least one fusion feature map in the embodiment of the present application;
图14为本申请实施例中获得至少一个目标特征图的示意图;FIG. 14 is a schematic diagram of obtaining at least one target feature map in an embodiment of the present application;
图15为本申请实施例获得目标车辆的可行驶区域的流程示意图;FIG. 15 is a schematic flow diagram of obtaining the drivable area of a target vehicle according to an embodiment of the present application;
图16为本申请实施例中包含可行驶区域的图像的示意图;Fig. 16 is a schematic diagram of an image including a drivable area in the embodiment of the present application;
图17为本申请实施例中可行驶区域识别装置的结构示意图;Fig. 17 is a schematic structural diagram of a drivable area recognition device in an embodiment of the present application;
图18为本申请实施例中电子设备的结构示意图。FIG. 18 is a schematic structural diagram of an electronic device in an embodiment of the present application.
具体实施方式Detailed ways
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本申请一部分实施例,并不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all of them. Based on the embodiments in this application, all other embodiments obtained by persons of ordinary skill in the art without making creative efforts belong to the scope of protection of this application.
为便于理解本申请实施例提供的技术方案,这里先对本申请实施例使用的一些关键名词进行解释:In order to facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application are first explained here:
可行驶区域:不考虑交通规则的情况下,除去车辆行驶道路上所有障碍物(包括护栏、车辆、行人以及其他障碍物)之外,允许车辆正常行驶的所有道路区域。Drivable area: Regardless of traffic rules, all road areas that allow vehicles to drive normally except for all obstacles (including guardrails, vehicles, pedestrians and other obstacles) on the road where the vehicle is driving.
损失函数:机器学习中用于度量和优化模型预估值和真实样本标注值之间距离的函数。Loss function: A function used in machine learning to measure and optimize the distance between the predicted value of the model and the labeled value of the real sample.
下面对本申请实施例的设计思想进行简要介绍:The design idea of the embodiment of the present application is briefly introduced below:
目前,检测车辆行驶过程中的场景信息是辅助驾驶系统和自动驾驶系统中的关键,场景信息的检测包括车道线、交通标识、道路标识及可行驶区域等,其中,可行驶区域检测在辅助驾驶系统和自动驾驶系统中发挥着重要作用,在车辆行驶过程中,为车辆提供可行驶路线,避开不可行驶区域,不可行驶区域主要包括道路上的静态障碍物,例如,锥桶、护栏等所在的区域,以及动态障碍物,例如,车辆、行人等所在的区域。At present, the detection of scene information during vehicle driving is the key to assisted driving systems and automatic driving systems. The detection of scene information includes lane lines, traffic signs, road signs, and drivable areas. The system and the automatic driving system play an important role. During the driving process of the vehicle, it provides the vehicle with a drivable route and avoids the undrivable area. The undrivable area mainly includes static obstacles on the road, such as cones, guardrails, etc. area, as well as the area where dynamic obstacles, such as vehicles and pedestrians, are located.
目前常用的可行驶区域识别方法包括:Currently commonly used drivable area identification methods include:
方式一:人工设计特征,获得可行驶区域,例如:传统的阈值法、聚类法、支持向量机、决策树等方法。Method 1: Artificially design features to obtain the drivable area, such as: traditional threshold method, clustering method, support vector machine, decision tree and other methods.
采用方法一时,由于需要人工设计特征,受专业理论和先验知识的限制,导致难以表征复杂交通环境的多变性及道路结构的多样性,只能在有限的特定环境下应用,具有局限性,并且效率低。When using method 1, due to the need for artificial design features, limited by professional theories and prior knowledge, it is difficult to represent the variability of complex traffic environments and the diversity of road structures, and it can only be applied in limited specific environments, which has limitations. And it is inefficient.
方式二:对采集的图片,使用复杂的编码器提取特征,获得四倍下采样特征图。例如,高分辨率网络(High-Resolution Net,HRNet),再将四倍下采样特征图进行放大处理,使得输出图片恢复到输入图片的尺寸,随后获得可行驶区域。Method 2: For the collected pictures, use a complex encoder to extract features and obtain a quadruple downsampled feature map. For example, the high-resolution network (High-Resolution Net, HRNet), and then enlarge the four times downsampled feature map, so that the output image can be restored to the size of the input image, and then the drivable area can be obtained.
采用方法二时,由于复杂的编码器的内存访问代价较大,因此图形处理器(graphics processingunit,GPU)的计算效率较低,限制了识别速度。When the second method is adopted, since the memory access cost of the complex encoder is high, the calculation efficiency of the graphics processing unit (GPU) is low, which limits the recognition speed.
另外,为了提升识别速度,直接采用放大处理恢复图片的尺寸,由于四倍下采样特征图的像素点较小,对象不能够观察清楚像素点,因此若直接对四倍下采样特征图进行放大处理,会造成识别不准确的情况。In addition, in order to improve the recognition speed, the size of the image is directly restored by zooming in. Since the pixels of the quadruple downsampling feature map are small, the object cannot observe the pixels clearly. Therefore, if the quadruple downsampling feature map is directly enlarged , resulting in inaccurate recognition.
例如,当四倍下采样特征图存在几个可行驶区域的像素点未被识别成可行驶区域的情况时,进行放大处理后,像素点会变大,对象能够观察到有一小块区域没有被识别成可行驶区域,导致出现空洞问题。For example, when the quadruple downsampling feature map has several pixels in the drivable area that are not recognized as drivable areas, after the zoom-in process, the pixels will become larger, and the object can observe that there is a small area that is not recognized as a drivable area. Identified as a drivable area, resulting in a hole problem.
有鉴于此,本申请实施例中,提出了一种可行驶区域识别方法、装置、电子设备及存储介质。在驾驶对象触发可行驶区域识别请求时,可以获取当前待识别的目标道路图像,然后按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图,进而针对各道路特征图,分别执行以下操作:基于一个道路特征图包含的各像素点各自的图像位置信息,结合一个道路特征图包含的图像通道信息,对道路特征图进行注意力加权处理,获得相应的注意力特征图,最后基于目标图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,并基于至少一个目标特征图,获得目标车辆的可行驶区域。这样,对不同尺度的道路特征图进行加权处理和融合操作,使目标特征图和目标道路图像尺寸保持一致,可以更好地平衡多尺度道路特征图的空间信息和语义信息,避免了空洞问题,提高了可行驶区域识别的结果准确率和识别效率。In view of this, in the embodiment of the present application, a drivable area recognition method, device, electronic device and storage medium are proposed. When the driving object triggers a drivable area recognition request, the current target road image to be recognized can be obtained, and then according to the preset candidate image scales, feature extraction is performed on the target road image to obtain the corresponding road feature map, and then for each The road feature map performs the following operations respectively: Based on the image position information of each pixel contained in a road feature map, combined with the image channel information contained in a road feature map, the road feature map is weighted by attention to obtain the corresponding attention The force feature map, and finally based on the target image scale, image scale adjustment is performed on each of the obtained attention feature maps to obtain at least one target feature map, and based on the at least one target feature map, the drivable area of the target vehicle is obtained. In this way, the weighted processing and fusion operations are performed on road feature maps of different scales to keep the size of the target feature map and the target road image consistent, which can better balance the spatial information and semantic information of the multi-scale road feature map, and avoid the hole problem. The result accuracy and recognition efficiency of drivable area recognition are improved.
以下结合说明书附图对本申请的优选实施例进行说明,应当理解,此处所描述的优选实施例仅用于说明和解释本申请,并不用于限定本申请,并且在不冲突的情况下,本申请实施例及实施例中的特征可以相互组合。The preferred embodiments of the application will be described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described here are only used to illustrate and explain the application, and are not used to limit the application, and in the absence of conflict, the application The embodiments and the features in the embodiments can be combined with each other.
参阅图1所示,为本申请实施例中可能的应用场景示意图。该应用场景图中,包括服务器110,以及终端设备120(包括终端设备1201、终端设备1202…终端设备120n)。Referring to FIG. 1 , it is a schematic diagram of a possible application scenario in the embodiment of the present application. In the application scenario diagram, a server 110 and terminal devices 120 (including terminal device 1201, terminal device 1202...terminal device 120n) are included.
服务器110可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络(Content Delivery Network,CDN)、以及大数据和人工智能平台等基础云计算服务的云服务器。终端设备120与服务器110可以通过有线或无线通信方式进行直接或间接地连接,本申请在此不做限制。The server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, Cloud servers for basic cloud computing services such as middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal device 120 and the server 110 may be connected directly or indirectly through wired or wireless communication, which is not limited in this application.
终端设备120可以是驾驶对象在目标车辆的行驶过程中,随身携带的手机、便携式电脑等,终端设备还可以是设置在目标车辆中的车载终端等具有一定计算能力的计算机设备。The terminal device 120 may be a mobile phone, a portable computer, etc. carried by the driving object during the driving of the target vehicle, and the terminal device may also be a computer device with certain computing capabilities such as a vehicle-mounted terminal installed in the target vehicle.
本申请实施例中,服务器110和终端设备120中的任意一方或全部,可以配置有已训练的目标识别模型,使得能够采用已训练的目标识别模型,对待识别的目标道路图像进行可行驶区域识别。In the embodiment of the present application, any one or all of the server 110 and the terminal device 120 can be configured with a trained target recognition model, so that the trained target recognition model can be used to identify the drivable area of the target road image to be recognized .
需要说明的是,本申请实施例中,设备上采用的识别模型可以是自身训练得到的,或者,可以是其他设备训练得到的。It should be noted that, in the embodiment of the present application, the recognition model adopted by the device may be obtained by its own training, or may be obtained by training of other devices.
以服务器110采用已训练的识别模型,对待识别的目标道路图像执行可行驶区域识别操作为例,服务器110采用的识别模型可以是自身训练得到的,或者,可以是其他设备完成识别模型的训练后,直接发送至服务器110的。Taking the server 110 using the trained recognition model to perform drivable area recognition operation on the target road image to be recognized as an example, the recognition model used by the server 110 can be obtained by its own training, or it can be obtained after other devices complete the training of the recognition model. , sent directly to the server 110.
以下的说明中,将以服务器实现目标识别模型的训练为例,对相关的训练过程进行详细说明。In the following description, the training of the object recognition model implemented by the server will be taken as an example to describe the related training process in detail.
另外,本申请实施例中,根据实际的处理需要,服务器训练得到目标识别模型可以是一个周期性的过程,可以周期性的重新生成训练样本,训练得到目标识别模型。In addition, in the embodiment of the present application, according to actual processing needs, the server training to obtain the target recognition model may be a periodic process, and the training samples may be periodically regenerated to train the target recognition model.
基于训练样本集,对待训练的识别模型进行迭代训练,获得目标识别模型;每个训练样本包括:样本道路的图像数据,参阅图2所示,为本申请实施例中目标识别模型一轮训练的流程示意图,具体包括:Based on the training sample set, the recognition model to be trained is iteratively trained to obtain the target recognition model; each training sample includes: the image data of the sample road, as shown in Figure 2, which is a round of training of the target recognition model in the embodiment of the present application Schematic diagram of the process, including:
步骤20:按照预设的各样本图像尺度,分别对选取的训练样本进行特征提取,获得相应的样本特征图。Step 20: Perform feature extraction on the selected training samples according to the preset image scales of each sample, and obtain corresponding sample feature maps.
其中,训练样本具有样本图像尺度。where the training samples have a sample image scale.
本申请实施例中,采用编码器对训练样本进行特征提取,按照预设的各样本图像尺度,分别对选取的训练样本进行特征提取,获得相应的样本特征图。In the embodiment of the present application, an encoder is used to perform feature extraction on training samples, and according to preset image scales of each sample, feature extraction is performed on selected training samples to obtain corresponding sample feature maps.
其中,预设的各样本图像尺度可以为:1/2、1/4、1/8、1/16,本申请实施例中对此并不进行限制。Wherein, the preset scales of each sample image may be: 1/2, 1/4, 1/8, 1/16, which is not limited in this embodiment of the present application.
例如,参阅图3所示,为本申请实施例中编码器的示意图,编码器主要由一个网络茎层和四个阶段层构成,网络茎层由3*3卷积层、批归一化层和FReLU层构成,获得图像尺寸为1/2的第一样本特征图,阶段层由高性能GPU模块(HG block)层和下采样(LDS Layer)层构成,下采样层为采用步长为2的2*2最大池化层,获得图像尺寸分别为1/4、1/8和1/16的各第一样本特征图,基于各第一样本特征图,采用HG block层,获得图像尺寸分别为1/2、1/4、1/8、1/16的样本特征图。For example, refer to Figure 3, which is a schematic diagram of the encoder in the embodiment of the present application. The encoder is mainly composed of a network stem layer and four stage layers. The network stem layer consists of a 3*3 convolutional layer and a batch normalization layer. Composed with FReLU layer, the first sample feature map with image size of 1/2 is obtained. The stage layer is composed of high-performance GPU module (HG block) layer and downsampling (LDS Layer) layer. The downsampling layer adopts a step size of 2's 2*2 maximum pooling layer to obtain the first sample feature maps with image sizes of 1/4, 1/8 and 1/16 respectively, based on each first sample feature map, using the HG block layer to obtain Sample feature maps with image sizes of 1/2, 1/4, 1/8, 1/16, respectively.
具体的,在HG Block层中,针对各样本图像尺度对应的第一样本特征图,分别执行以下操作:按照预设的各样本图像感受野,分别对一个样本图像尺度对应的第一样本特征图进行第二特征提取,获得相应的第二样本特征图,并对获得的各第二样本特征图进行叠加,获得叠加样本特征图,基于第一样本特征图的图像通道数,对叠加特征图进行图像通道调整,获得该样本图像尺度对应的样本特征图。Specifically, in the HG Block layer, for the first sample feature map corresponding to each sample image scale, the following operations are respectively performed: according to the preset receptive field of each sample image, the first sample corresponding to a sample image scale The second feature extraction is performed on the feature map to obtain the corresponding second sample feature map, and the obtained second sample feature maps are superimposed to obtain the superimposed sample feature map. Based on the number of image channels of the first sample feature map, the superimposed The feature map performs image channel adjustment to obtain the sample feature map corresponding to the sample image scale.
其中,每个样本图像感受野表征:相应第二样本特征图包含的各像素点映射至训练样本上的区域大小,各样本图像感受野可以为:1x1、3x 3、5x 5、7x 7、9x 9、11x 11,本申请实施例中对此并不进行限制。Among them, the representation of the receptive field of each sample image: each pixel contained in the corresponding second sample feature map is mapped to the size of the area on the training sample, and the receptive field of each sample image can be: 1x1, 3x 3, 5x 5, 7x 7, 9x 9. 11 x 11, which is not limited in the embodiment of the present application.
例如,参阅图4所示,为本申请实施例中高性能GPU模块层的示意图,第一样本特征图先通过串联的3*3卷积层,获得不同图像感受野的第二样本特征图包括:感受野为3x 3的第二样本特征图、感受野为5x 5的第二样本特征图、感受野为7x 7的第二样本特征图、感受野为9x 9的第二样本特征图、感受野为11x 11的第二样本特征图,并对获得的各第二样本特征图进行叠加,获得叠加样本特征图,采用1*1卷积对叠加样本特征图进行图像通道调整,获得样本特征图,使样本特征图的图像通道数与第一样本特征图的图像通道数相同。For example, refer to FIG. 4, which is a schematic diagram of the high-performance GPU module layer in the embodiment of the present application. The first sample feature map first passes through a series of 3*3 convolutional layers to obtain the second sample feature map of different image receptive fields. : The second sample feature map with a receptive field of 3x3, the second sample feature map with a receptive field of 5x5, the second sample feature map with a receptive field of 7x7, the second sample feature map with a receptive field of 9x9, and the receptive The field is the second sample feature map of 11x11, and superimpose the obtained second sample feature maps to obtain the superimposed sample feature map, and use 1*1 convolution to adjust the image channel of the superimposed sample feature map to obtain the sample feature map , so that the number of image channels of the sample feature map is the same as that of the first sample feature map.
这样,由于内存访问代价对计算效率的影响远大于模型的计算量和参数量,特别是对于模型中大量使用的卷积层,只有当输入和输出的图像通道数相等时,内存访问代价最小,因此,采用高性能GPU模块层,采用输入和输出的图像通道数相同的设计,可以最小化内存访问代价,在提高计算效率的同时,通过分层卷积重叠提升网络的多特征提取能力。In this way, since the impact of memory access cost on computational efficiency is much greater than the amount of computation and parameters of the model, especially for convolutional layers that are heavily used in the model, only when the number of input and output image channels is equal, the memory access cost is the smallest. Therefore, adopting a high-performance GPU module layer and adopting the design with the same number of input and output image channels can minimize the cost of memory access, improve computing efficiency, and improve the multi-feature extraction capability of the network through layered convolution overlap.
可选的,在最后一个阶段层包含的高性能GPU模块层中添加残差连接和挤压激励,如图4所示,虚线部分为残差连接,这样,由于残差连接在加快模型收敛、避免梯度消失等方面表现优异,但是残差连接所带来的元素级(element-wise)加法操作不利于GPU设备计算,会影响模型的速度。挤压激励是一种通道注意力机制,可以有效提升模型的精度,但是同样会带来较大的延时。为了平衡模型的精度和速度,编码器仅在最后一个阶段层中使用残差连接和挤压激励。Optionally, add residual connection and squeeze excitation to the high-performance GPU module layer included in the last stage layer, as shown in Figure 4, the dotted line part is the residual connection, in this way, since the residual connection speeds up the model convergence, Excellent performance in avoiding gradient disappearance, etc., but the element-wise addition operation brought by the residual connection is not conducive to GPU device calculation and will affect the speed of the model. Squeeze excitation is a channel attention mechanism, which can effectively improve the accuracy of the model, but it will also bring a large delay. To balance the accuracy and speed of the model, the encoder only uses residual connections and squeeze excitations in the last stage layer.
步骤21:针对各样本特征图,分别执行以下操作:基于一个样本特征图包含的各像素点各自的样本图像位置信息,结合一个样本特征图包含的样本图像通道信息,对样本特征图进行注意力加权处理,获得相应的样本注意力特征图。Step 21: For each sample feature map, perform the following operations: Based on the sample image position information of each pixel contained in a sample feature map, combined with the sample image channel information contained in a sample feature map, pay attention to the sample feature map Weighted processing to obtain the corresponding sample attention feature map.
本申请实施例中,针对各样本特征图,分别执行以下操作:采用空间注意力模块,基于一个样本特征图包含的各像素点各自的样本图像位置信息,确定该样本特征图包含的各像素点各自的样本位置注意力权重,采用通道注意力模块,基于该样本特征图包含的样本图像通道信息,确定该样本特征图包含的各图像通道各自的样本通道注意力权重,基于获得的各样本位置注意力权重和样本通道注意力权重,对该样本特征图进行加权,获得该样本特征图对应的样本注意力特征图。In the embodiment of the present application, for each sample feature map, the following operations are respectively performed: using a spatial attention module, based on the respective sample image position information of each pixel contained in a sample feature map, determine each pixel contained in the sample feature map Attention weights of respective sample positions, using the channel attention module, based on the sample image channel information contained in the sample feature map, determine the respective sample channel attention weights of each image channel contained in the sample feature map, based on the obtained sample position Attention weight and sample channel attention weight weight the sample feature map to obtain the sample attention feature map corresponding to the sample feature map.
其中,样本位置注意力权重表征在识别可行驶区域时,相应像素点的样本图像位置的重要程度,样本通道注意力权重表征在识别可行驶区域时,相应样本图像通道的重要程度。Among them, the sample position attention weight represents the importance of the sample image position of the corresponding pixel when identifying the drivable area, and the sample channel attention weight represents the importance of the corresponding sample image channel when identifying the drivable area.
具体的,参阅图5所示,为本申请实施例中空间注意力模块和通道注意力模块的示意图,空间注意力模块是在空间位置维度对样本特征图进行平均池化(Mean Pooling)和最大池化(Max Pooling)操作,获得空间平均池化样本特征图和空间最大池化样本特征图,然后将空间平均池化样本特征图和空间最大池化样本特征图进行拼接,获得空间池化拼接样本特征图,最后通过卷积层和sigmoid函数激活,获得该样本特征图包含的各样本像素点各自的样本位置注意力权重;通道注意力模块是在通道维度对样本特征图进行Mean Pooling和Max Pooling操作,获得通道平均池化样本特征图和通道最大池化样本特征图,然后将通道平均池化样本特征图和通道最大池化样本特征图进行拼接,获得通道池化拼接样本特征图,最后通过卷积层和sigmoid函数激活,获得该样本特征图包含的各样本图像通道各自的样本通道注意力权重。Specifically, refer to FIG. 5, which is a schematic diagram of the spatial attention module and the channel attention module in the embodiment of the present application. The spatial attention module performs mean pooling (Mean Pooling) and maximum The pooling (Max Pooling) operation obtains the spatial average pooling sample feature map and the spatial maximum pooling sample feature map, and then stitches the spatial average pooling sample feature map and the spatial maximum pooling sample feature map to obtain the spatial pooling splicing The sample feature map is finally activated by the convolutional layer and the sigmoid function to obtain the respective sample position attention weights of each sample pixel contained in the sample feature map; the channel attention module performs Mean Pooling and Max on the sample feature map in the channel dimension The Pooling operation obtains the channel average pooling sample feature map and the channel maximum pooling sample feature map, and then stitches the channel average pooling sample feature map and the channel maximum pooling sample feature map to obtain the channel pooling splicing sample feature map, and finally Through the activation of the convolutional layer and the sigmoid function, the respective sample channel attention weights of the sample image channels included in the sample feature map are obtained.
这样,采用空间注意力模块和通道注意力模块,获得各样本注意力特征图,能够让识别模型更加关注样本特征图中的重要信息,抑制无用信息,提高识别模型的准确度。In this way, using the spatial attention module and the channel attention module to obtain the attention feature map of each sample can make the recognition model pay more attention to the important information in the sample feature map, suppress useless information, and improve the accuracy of the recognition model.
步骤22:基于样本图像尺度,分别对获得的各样本注意力特征图进行图像尺度调整,获得至少一个目标样本特征图。Step 22: Based on the sample image scale, perform image scale adjustment on the obtained attention feature maps of each sample to obtain at least one target sample feature map.
本申请实施例中,基于获得的各样本注意力特征图各自的图像尺度大小,对各样本注意力特征图进行融合更新,获得至少一个样本融合特征图,基于预设的样本图像通道数,分别对至少一个样本融合特征图进行图像通道调整,获得相应的样本中间特征图,基于样本图像尺度,对获得的至少一个样本中间特征图进行图像尺度调整,获得相应的样本目标特征图。In the embodiment of the present application, based on the respective image scales of the acquired attention feature maps of each sample, the attention feature maps of each sample are fused and updated to obtain at least one sample fusion feature map. Based on the preset number of sample image channels, respectively Image channel adjustment is performed on at least one sample fusion feature map to obtain a corresponding sample intermediate feature map, and based on the sample image scale, image scale adjustment is performed on the obtained at least one sample intermediate feature map to obtain a corresponding sample target feature map.
例如,采用卷积网络,基于预设的样本图像通道数,分别对至少一个样本融合特征图进行图像通道调整,获得相应的样本中间特征图,采用双线性插值,基于样本图像尺度,对获得的至少一个样本中间特征图进行图像尺度调整,获得相应的样本目标特征图。For example, using a convolutional network, based on the preset number of sample image channels, adjust the image channel of at least one sample fusion feature map to obtain the corresponding sample intermediate feature map, using bilinear interpolation, based on the sample image scale, to obtain Image scaling is performed on at least one sample intermediate feature map to obtain the corresponding sample target feature map.
这样,获得样本图像通道数和样本图像尺度都相同的至少一个样本目标特征图,整合了不同样本注意力特征图在通道和尺度特征上的差异,提高了目标识别模型的准确度。In this way, at least one sample target feature map with the same sample image channel number and sample image scale is obtained, and the differences in channel and scale features of different sample attention feature maps are integrated to improve the accuracy of the target recognition model.
具体的,对各样本注意力特征图进行融合更新有多种方法。Specifically, there are many ways to fuse and update the attention feature maps of each sample.
第一种融合更新方法:按照获得的各样本注意力特征图各自的图像尺度大小,将各样本注意力特征图进行排序,获得样本排序结果,并基于样本排序结果,依次读取每两个相邻的样本注意力特征图进行融合,直到读取完毕为止,其中,一次融合过程包括:读取相邻的两个样本注意力特征图,并对该两个样本注意力特征图进行融合,获得样本融合特征图,保存样本融合特征图,以及将样本融合特征图作为新的样本注意力特征图,代替两个样本注意力特征图,加入样本排序结果中。The first fusion update method: sort the attention feature maps of each sample according to the respective image scales of the obtained attention feature maps of each sample, obtain the sample sorting results, and read every two phases in sequence based on the sample sorting results. Neighboring sample attention feature maps are fused until the reading is completed. A fusion process includes: reading two adjacent sample attention feature maps, and fusing the two sample attention feature maps to obtain Sample fusion feature map, save the sample fusion feature map, and use the sample fusion feature map as a new sample attention feature map, replace the two sample attention feature maps, and add it to the sample sorting result.
例如,假设各样本注意力特征图包括样本注意力特征图A1、样本注意力特征图A2、样本注意力特征图A3、样本注意力特征图A4,样本注意力特征图A1的图像尺度大小为1/2、样本注意力特征图A2的图像尺度大小为1/4、样本注意力特征图A3的图像尺度大小为1/8、样本注意力特征图A4的图像尺度大小为1/16,则将各样本注意力特征图按照图像尺度从小到大进行排序,获得的样本排序结果为A1、A2、A3、A4,读取A1和A2,并对A1和A2进行融合,获得样本融合特征图B1,保存B1,以及将B1作为新的样本注意力特征图,代替A1和A2,加入样本排序结果中,此时样本排序结果为B1、A3、A4,读取B1和A3,并对B1和A3进行融合,获得样本融合特征图B2,保存B2,以及将B2作为新的样本注意力特征图,代替B1和A3,加入样本排序结果中,此时样本排序结果为B2、A4,读取B2和A4,并对B2和A4进行融合,获得样本融合特征图B3,保存B3,最终获得的至少一个样本融合特征图包括B1、B2和B3。For example, suppose each sample attention feature map includes sample attention feature map A1, sample attention feature map A2, sample attention feature map A3, sample attention feature map A4, and the image scale of sample attention feature map A1 is 1 /2. The image scale of the sample attention feature map A2 is 1/4, the image scale of the sample attention feature map A3 is 1/8, and the image scale of the sample attention feature map A4 is 1/16, then the The attention feature maps of each sample are sorted according to the image scale from small to large, and the obtained sample sorting results are A1, A2, A3, A4, read A1 and A2, and fuse A1 and A2 to obtain the sample fusion feature map B1, Save B1, and use B1 as a new sample attention feature map, replace A1 and A2, and add it to the sample sorting result. At this time, the sample sorting result is B1, A3, A4, read B1 and A3, and perform B1 and A3 Fusion, obtain the sample fusion feature map B2, save B2, and use B2 as a new sample attention feature map, replace B1 and A3, and add it to the sample sorting result. At this time, the sample sorting result is B2, A4, read B2 and A4 , and fuse B2 and A4 to obtain a sample fusion feature map B3, save B3, and finally obtain at least one sample fusion feature map including B1, B2 and B3.
进一步地,不同样本图像尺度对应的样本注意力特征图之间的融合更新之前,将小样本图像尺度对应的样本特征图上采样至与大样本图像尺度对应的样本特征图相同的样本图像尺度,然后再通过空间注意力模块和通道注意力模块,获得小样本图像尺度对应的样本注意力特征图。Further, before the fusion update between the sample attention feature maps corresponding to different sample image scales, the sample feature map corresponding to the small sample image scale is up-sampled to the same sample image scale as the sample feature map corresponding to the large sample image scale, Then through the spatial attention module and the channel attention module, the sample attention feature map corresponding to the small sample image scale is obtained.
第二种融合更新方法:按照获得的各样本注意力特征图各自的图像尺度大小,将各样本注意力特征图进行排序,获得样本排序结果,并基于样本排序结果,依次读取每两个相邻的样本注意力特征图进行融合,直到读取完毕为止,其中,一次融合过程包括:读取相邻的两个样本注意力特征图,并对该两个样本注意力特征图进行融合,获得样本融合特征图,保存样本融合特征图。The second fusion update method: sort the attention feature maps of each sample according to the respective image scales of the obtained attention feature maps of each sample, obtain the sample sorting results, and read every two phases in sequence based on the sample sorting results. Neighboring sample attention feature maps are fused until the reading is completed. A fusion process includes: reading two adjacent sample attention feature maps, and fusing the two sample attention feature maps to obtain Sample fusion feature map, save the sample fusion feature map.
例如,假设各样本注意力特征图包括样本注意力特征图A1、样本注意力特征图A2、样本注意力特征图A3、样本注意力特征图A4,样本注意力特征图A1的图像尺度大小为1/2、样本注意力特征图A2的图像尺度大小为1/4、样本注意力特征图A3的图像尺度大小为1/8、样本注意力特征图A4的图像尺度大小为1/16,则将各样本注意力特征图按照图像尺度从小到大进行排序,获得的样本排序结果为A1、A2、A3、A4,读取A1和A2,并对A1和A2进行融合,获得样本融合特征图B4,保存B4,读取A2和A3,并对A2和A3进行融合,获得样本融合特征图B5,保存B5,读取A3和A4,并对A3和A4进行融合,获得样本融合特征图B6,保存B6,最终获得的至少一个样本融合特征图包括B4、B5和B6。For example, suppose each sample attention feature map includes sample attention feature map A1, sample attention feature map A2, sample attention feature map A3, sample attention feature map A4, and the image scale of sample attention feature map A1 is 1 /2. The image scale of the sample attention feature map A2 is 1/4, the image scale of the sample attention feature map A3 is 1/8, and the image scale of the sample attention feature map A4 is 1/16, then the The attention feature map of each sample is sorted according to the image scale from small to large, and the obtained sample sorting results are A1, A2, A3, A4, read A1 and A2, and fuse A1 and A2 to obtain the sample fusion feature map B4, Save B4, read A2 and A3, and fuse A2 and A3, obtain the sample fusion feature map B5, save B5, read A3 and A4, and fuse A3 and A4, obtain the sample fusion feature map B6, save B6 , the finally obtained at least one sample fusion feature map includes B4, B5 and B6.
第三种融合更新方法:将各样本注意力特征图进行融合,获得一个样本融合特征图。The third fusion update method: fuse the attention feature maps of each sample to obtain a sample fusion feature map.
例如,假设各样本注意力特征图包括样本注意力特征图A1、样本注意力特征图A2、样本注意力特征图A3、样本注意力特征图A4,则将A1、A2、A3、A4进行融合,获得一个样本融合特征图B7。For example, assuming that each sample attention feature map includes sample attention feature map A1, sample attention feature map A2, sample attention feature map A3, and sample attention feature map A4, then A1, A2, A3, and A4 are fused, Obtain a sample fusion feature map B7.
步骤23:基于至少一个目标样本特征图,获得训练样本的可行驶区域预测结果,并基于可行驶区域预测结果对应的损失值进行调参。Step 23: Obtain the drivable area prediction result of the training sample based on at least one target sample feature map, and perform parameter adjustment based on the loss value corresponding to the drivable area prediction result.
本申请实施例中,基于至少一个目标样本特征图,获得样本结果特征图,将样本结果特征图中包含的各样本像素点的值与预设的样本可行驶区域标签值进行比较,确定样本结果特征图中属于可行驶区域的各样本目标像素点,将训练样本中,与各样本目标像素点各自的图像位置对应的各样本像素点进行标示,获得训练样本的可行驶区域预测结果,并基于可行驶区域预测结果、训练样本对应的真实标签和难易样本权重,计算损失值,以及基于损失值,调整识别模型的网络参数。In the embodiment of the present application, the sample result feature map is obtained based on at least one target sample feature map, and the value of each sample pixel contained in the sample result feature map is compared with the preset sample drivable area label value to determine the sample result Each sample target pixel point belonging to the drivable area in the feature map is marked with each sample pixel point corresponding to the respective image position of each sample target pixel point in the training sample to obtain the drivable area prediction result of the training sample, and based on The prediction result of the drivable area, the real label corresponding to the training sample and the weight of the difficult sample, calculate the loss value, and adjust the network parameters of the recognition model based on the loss value.
本申请实施例中,获得样本结果特征图,包含但不限于以下两种情况:In the embodiment of this application, the sample result feature map is obtained, including but not limited to the following two situations:
情况1:至少一个目标样本特征图中只包括一个目标样本特征图,将一个目标样本特征图作为样本结果特征图。Case 1: At least one target sample feature map includes only one target sample feature map, and one target sample feature map is used as the sample result feature map.
情况2:至少一个目标样本特征图包括多个目标样本特征图,对多个目标样本特征图进行图像融合,获得样本结果特征图。Case 2: At least one target sample feature map includes multiple target sample feature maps, image fusion is performed on the multiple target sample feature maps to obtain a sample result feature map.
其中,预设的样本可行驶区域标签值可以为1,本申请实施例对此并不进行限制,模型训练的目标函数为loss=0.9*LFocalLoss+0.1*LLovaszLoss,LFocalLoss为根据识别的难易程度施加对应的权重损失,即为易识别的样本添加较小的权重,为难识别的样本添加较大的权重,LLovaszLoss为基于子模损失的凸Lovasz扩展,对语义分割任务的均交并比(meanintersection over union,Mean IoU)进行优化。Among them, the preset sample drivable area label value can be 1, which is not limited in the embodiment of the present application. The objective function of model training is loss=0.9*L FocalLoss +0.1*L LovaszLoss , and L FocalLoss is based on the identified The difficulty level applies the corresponding weight loss, that is, adding a smaller weight to the easy-to-recognize samples and adding a larger weight to the difficult-to-recognize samples. L LovaszLoss is a convex Lovasz extension based on the sub-module loss, which is uniform for semantic segmentation tasks. And optimize (meanintersection over union, Mean IoU).
例如,采用sigmoid函数对样本结果特征图进行激活,将样本结果特征图中包含的各样本像素点的值与预设的样本可行驶区域标签值1进行比较,确定样本结果特征图中属于可行驶区域的各样本目标像素点,将训练样本中,与各样本目标像素点各自的图像位置对应的各样本像素点进行标示,获得训练样本的可行驶区域预测结果。For example, use the sigmoid function to activate the sample result feature map, compare the value of each sample pixel contained in the sample result feature map with the preset sample drivable area label value 1, and determine that the sample result feature map belongs to the drivable area. For each sample target pixel point in the area, each sample pixel point corresponding to the image position of each sample target pixel point in the training sample is marked, and the drivable area prediction result of the training sample is obtained.
下面结合附图,对基于已训练的目标识别模型进行应用的过程进行说明:The following describes the application process based on the trained target recognition model with reference to the accompanying drawings:
参阅图6所示,其为本申请实施例中可行驶区域识别方法的流程示意图,下面结合附图6,对具体操作进行详细说明:Refer to Figure 6, which is a schematic flow chart of the drivable area identification method in the embodiment of the present application. The specific operations will be described in detail below in conjunction with Figure 6:
步骤60:响应于驾驶对象触发的可行驶区域识别请求,获取当前待识别的目标道路图像。Step 60: Responding to the drivable area identification request triggered by the driving object, acquire the current target road image to be identified.
其中,目标道路图像具有目标图像尺度。Wherein, the target road image has a target image scale.
本申请实施例中,驾驶对象在终端设备上的进行一定的操作之后,则可以触发终端设备向服务器发送可行驶区域识别请求,服务器响应于驾驶对象触发的可行驶区域识别请求,调用拍摄设备,获取当前待识别的目标道路图像。In this embodiment of the application, after the driving object performs certain operations on the terminal device, the terminal device can be triggered to send a drivable area identification request to the server, and the server calls the shooting device in response to the drivable area identification request triggered by the driving object. Obtain the current target road image to be recognized.
步骤61:按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图。Step 61: According to the preset scales of each candidate image, feature extraction is performed on the target road image to obtain a corresponding road feature map.
本申请实施例中,获得目标道路图像之后,按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得各候选图像尺度各自对应的道路特征图。In the embodiment of the present application, after the target road image is obtained, feature extraction is performed on the target road image according to the preset candidate image scales, and road feature maps corresponding to each candidate image scale are obtained.
其中,预设的各候选图像尺度可以为:1/2、1/4、1/8、1/16,本申请实施例中对此并不进行限制。The preset candidate image scales may be: 1/2, 1/4, 1/8, 1/16, which is not limited in this embodiment of the present application.
具体的,在获得一个候选图像尺度对应的道路特征图时,服务器具体执行以下操作。参阅图7所示,其为本申请实施例中获得道路特征图的流程示意图,下面结合附图7,对具体执行的操作进行详细说明:Specifically, when obtaining a road feature map corresponding to a candidate image scale, the server specifically performs the following operations. Referring to Fig. 7, it is a schematic flow chart of obtaining a road feature map in the embodiment of the present application. In conjunction with Fig. 7, the specific operations are described in detail below:
步骤610:基于一个候选图像尺度,对目标道路图像进行第一特征提取,获得第一特征图。Step 610: Based on a candidate image scale, perform first feature extraction on the target road image to obtain a first feature map.
其中,第一特征图具有第一图像通道数。Wherein, the first feature map has a first number of image channels.
本申请实施例中,基于一个候选图像尺度,采用相应的下采样操作,对目标道路图像进行第一特征提取,获得第一特征图,第一特征图的图像尺度为候选图像尺度。In the embodiment of the present application, based on a candidate image scale, a corresponding downsampling operation is used to perform first feature extraction on the target road image to obtain a first feature map, and the image scale of the first feature map is the candidate image scale.
例如,假设一个候选图像尺度为1/2,则对目标道路图像进行第一特征提取,获得图像尺度为1/2的第一特征图。For example, assuming that the scale of a candidate image is 1/2, the first feature extraction is performed on the target road image to obtain the first feature map with an image scale of 1/2.
步骤611:按照预设的各候选图像感受野,分别对第一特征图进行第二特征提取,获得相应的第二特征图,并对获得的各第二特征图进行叠加,获得叠加特征图。Step 611: Perform second feature extraction on the first feature map according to the preset receptive field of each candidate image to obtain a corresponding second feature map, and superimpose the obtained second feature maps to obtain a superimposed feature map.
其中,每个候选图像感受野表征:相应第二特征图包含的各像素点映射至目标道路图像上的区域大小,各候选图像感受野可以为:1x1、3x 3、5x 5、7x 7、9x 9、11x 11,本申请实施例中对此并不进行限制。Among them, the receptive field of each candidate image is characterized by: the size of the area where each pixel contained in the corresponding second feature map is mapped to the target road image, and the receptive field of each candidate image can be: 1x1, 3x 3, 5x 5, 7x 7, 9x 9. 11 x 11, which is not limited in the embodiment of the present application.
本申请实施例中,获得第一特征图之后,针对预设的各候选图像感受野,分别执行以下操作:基于一个候选图像感受野,对第一特征图进行第二特征提取,获得相应的第二特征图。并对获得的各第二特征图进行叠加,获得叠加特征图。In the embodiment of the present application, after the first feature map is obtained, the following operations are respectively performed for each preset candidate image receptive field: Based on a candidate image receptive field, the second feature extraction is performed on the first feature map to obtain the corresponding first Two feature maps. and superimposing the obtained second feature maps to obtain superimposed feature maps.
例如,参阅图8所示,为本申请实施例中获得叠加特征图的示意图,假设各候选图像感受野为:1x1、3x 3、5x 5、7x 7、9x 9、11x 11,采用串联的3*3卷积层,获得不同图像感受野的第二特征图包括:感受野为3x 3的第二特征图、感受野为5x 5的第二特征图、感受野为7x 7的第二特征图、感受野为9x 9的第二特征图、感受野为11x 11的第二特征图,并对获得的各第二样本特征图进行叠加,获得叠加特征图。For example, refer to Figure 8, which is a schematic diagram of obtaining a superimposed feature map in the embodiment of the present application, assuming that the receptive fields of each candidate image are: 1x1, 3x 3, 5x 5, 7x 7, 9x 9, 11x 11, using a series of 3 *3 convolutional layers, to obtain the second feature map of different image receptive fields, including: the second feature map with a receptive field of 3x 3, the second feature map with a receptive field of 5x 5, and the second feature map with a receptive field of 7x 7 , a second feature map with a receptive field of 9x9, and a second feature map with a receptive field of 11x11, and superimpose the obtained second sample feature maps to obtain a superimposed feature map.
步骤612:基于第一图像通道数,对叠加特征图进行图像通道调整,获得一个候选图像尺度对应的道路特征图。Step 612: Based on the number of first image channels, perform image channel adjustment on the superimposed feature map to obtain a road feature map corresponding to a candidate image scale.
本申请实施例中,获得叠加特征图之后,基于第一图像通道数,采用1*1卷积对叠加特征图进行图像通道调整,获得道路特征图,道路特征图的图像通道数为第一图像通道数。In the embodiment of the present application, after obtaining the superimposed feature map, based on the first image channel number, 1*1 convolution is used to adjust the image channel of the superimposed feature map to obtain the road feature map, and the image channel number of the road feature map is the first image number of channels.
这样,由于内存访问代价对计算效率的影响远大于模型的计算量和参数量,特别是对于卷积操作,只有当输入和输出的图像通道数相等时,内存访问代价最小,因此,采用道路特征图和第一特征图的图像通道数相同的设计,可以最小化内存访问代价,在提高计算效率的同时,通过分层卷积叠加,获得目标道路图像的多特征。In this way, since the impact of memory access cost on computational efficiency is much greater than the amount of computation and parameters of the model, especially for convolution operations, only when the number of input and output image channels is equal, the memory access cost is the smallest. Therefore, the road feature The design with the same number of image channels in the first feature map and the first feature map can minimize the memory access cost, and while improving the computational efficiency, multiple features of the target road image can be obtained through layered convolution and superposition.
可选的,基于第一图像通道数,对图像尺度最小的第一特征图对应的叠加特征图进行图像通道调整,获得最小候选图像尺度对应的道路特征图时,添加残差连接和挤压激励,如图8所示,虚线部分为残差连接,这样,由于残差连接在避免梯度消失等方面表现优异,但是残差连接所带来的元素级(element-wise)加法操作不利于GPU设备计算,会影响可行驶区域识别的速度。挤压激励是一种通道注意力机制,可以有效提升可行驶区域识别的精度,但是同样会带来较大的延时。为了平衡可行驶区域识别的精度和速度,仅在获得最小候选图像尺度对应的道路特征图时,使用残差连接和挤压激励。Optionally, based on the number of first image channels, image channel adjustment is performed on the superimposed feature map corresponding to the first feature map with the smallest image scale, and when obtaining the road feature map corresponding to the smallest candidate image scale, add residual connection and squeeze excitation , as shown in Figure 8, the dotted line part is the residual connection, in this way, because the residual connection is excellent in avoiding gradient disappearance, etc., but the element-wise addition operation brought by the residual connection is not conducive to GPU equipment calculation, which affects the speed of drivable area recognition. Squeeze incentive is a channel attention mechanism, which can effectively improve the accuracy of drivable area recognition, but it will also bring a large delay. In order to balance the accuracy and speed of drivable area recognition, residual connection and squeeze excitation are only used when obtaining the road feature map corresponding to the smallest candidate image scale.
另外,值得说明的是,本申请实施例中,可以先通过3*3卷积层、批归一化层和FReLU层,获得图像尺度为1/2的第一特征图,并获得图像尺度为1/2的第一特征图对应的道路特征图,再采用步长为2的2*2最大池化层,对图像尺度为1/2的第一特征图对应的道路特征图进行特征提取,获得图像尺度为1/4的第一特征图,并获得图像尺度为1/4的第一特征图对应的道路特征图,再采用步长为2的2*2最大池化层,对图像尺度为1/4的第一特征图对应的道路特征图进行特征提取,获得图像尺度为1/8的第一特征图,并获得图像尺度为1/8的第一特征图对应的道路特征图,再采用步长为2的2*2最大池化层,对图像尺度为1/8的第一特征图对应的道路特征图进行特征提取,获得图像尺度为1/16的第一特征图,并获得图像尺度为1/16的第一特征图对应的道路特征图。In addition, it is worth noting that in the embodiment of the present application, the first feature map with an image scale of 1/2 can be obtained through the 3*3 convolution layer, batch normalization layer and FReLU layer, and the image scale is 1/2 of the road feature map corresponding to the first feature map, and then use a 2*2 maximum pooling layer with a step size of 2 to perform feature extraction on the road feature map corresponding to the first feature map with an image scale of 1/2, Obtain the first feature map with an image scale of 1/4, and obtain the road feature map corresponding to the first feature map with an image scale of 1/4, and then use a 2*2 maximum pooling layer with a step size of 2 to Perform feature extraction for the road feature map corresponding to the first feature map of 1/4, obtain the first feature map with an image scale of 1/8, and obtain the road feature map corresponding to the first feature map with an image scale of 1/8, Then use a 2*2 maximum pooling layer with a step size of 2 to perform feature extraction on the road feature map corresponding to the first feature map with an image scale of 1/8 to obtain the first feature map with an image scale of 1/16, and The road feature map corresponding to the first feature map whose image scale is 1/16 is obtained.
步骤62:针对各道路特征图,分别执行以下操作:基于一个道路特征图包含的各像素点各自的图像位置信息,结合一个道路特征图包含的图像通道信息,对道路特征图进行注意力加权处理,获得相应的注意力特征图。Step 62: For each road feature map, perform the following operations respectively: based on the respective image position information of each pixel contained in a road feature map, combined with the image channel information contained in a road feature map, perform attention weighting processing on the road feature map , to obtain the corresponding attention feature map.
具体的,在执行步骤62时,服务器具体执行以下操作。参阅图9所示,其为本申请实施例中获得一个道路特征图对应的注意力特征图的流程示意图,下面结合附图9,对具体执行的操作进行详细说明:Specifically, when executing step 62, the server specifically performs the following operations. Referring to Figure 9, it is a schematic flow chart of obtaining an attention feature map corresponding to a road feature map in the embodiment of the present application. The specific operations performed will be described in detail below in conjunction with Figure 9:
步骤620:基于一个道路特征图包含的各像素点各自的图像位置信息,确定一个道路特征图包含的各像素点各自的位置注意力权重。Step 620: Based on the respective image position information of each pixel contained in a road feature map, determine the respective location attention weights of each pixel contained in a road feature map.
其中,位置注意力权重表征在识别可行驶区域时,相应像素点的图像位置的重要程度。Among them, the position attention weight represents the importance of the image position of the corresponding pixel when identifying the drivable area.
本申请实施例中,基于一个道路特征图包含的各像素点各自的图像位置信息,在空间位置维度对道路特征图进行平均池化和最大池化操作,获得空间平均池化特征图和空间最大池化特征图,然后将空间平均池化特征图和空间最大池化特征图进行拼接,获得空间池化拼接特征图,最后通过卷积层和sigmoid函数激活,获得该道路特征图包含的各像素点各自的位置注意力权重。In the embodiment of the present application, based on the respective image position information of each pixel contained in a road feature map, average pooling and maximum pooling operations are performed on the road feature map in the spatial position dimension to obtain the spatial average pooling feature map and the spatial maximum Pooling the feature map, and then splicing the spatial average pooling feature map and the spatial maximum pooling feature map to obtain the spatial pooling splicing feature map, and finally through the convolution layer and sigmoid function activation to obtain each pixel contained in the road feature map Point the respective positional attention weights.
步骤621:基于一个道路特征图包含的图像通道信息,确定一个道路特征图包含的各图像通道各自的通道注意力权重。Step 621: Based on the image channel information contained in a road feature map, determine the respective channel attention weights of each image channel contained in a road feature map.
其中,通道注意力权重表征在识别可行驶区域时,相应图像通道的重要程度。Among them, the channel attention weight represents the importance of the corresponding image channel when identifying the drivable area.
本申请实施例中,基于一个道路特征图包含的图像通道信息,在通道维度对道路特征图进行平均池化和最大池化操作,获得通道平均池化特征图和通道最大池化特征图,然后将通道平均池化特征图和通道最大池化特征图进行拼接,获得通道池化拼接特征图,最后通过卷积层和sigmoid函数激活,获得该道路特征图包含的各图像通道各自的通道注意力权重。In the embodiment of the present application, based on the image channel information contained in a road feature map, the average pooling and maximum pooling operations are performed on the road feature map in the channel dimension to obtain the channel average pooling feature map and the channel maximum pooling feature map, and then The channel average pooling feature map and the channel maximum pooling feature map are spliced to obtain the channel pooling stitching feature map, and finally the convolution layer and the sigmoid function are activated to obtain the channel attention of each image channel contained in the road feature map Weights.
步骤622:基于获得的各位置注意力权重和通道注意力权重,对一个道路特征图进行加权,获得一个道路特征图对应的注意力特征图。Step 622: Based on the obtained attention weights of each position and channel attention weight, weight a road feature map to obtain an attention feature map corresponding to a road feature map.
本申请实施例中,获得的各位置注意力权重和通道注意力权重之后,在空间位置维度,将一个道路特征图与其各位置注意力权重相乘,获得位置注意力特征图,在通道维度,将该道路特征图与其各通道注意力权重相乘,获得通道注意力特征图,最后将位置注意力特征图和通道注意力特征图相加,获得该道路特征图对应的注意力特征图。In the embodiment of the present application, after obtaining the position attention weights and channel attention weights, in the spatial position dimension, a road feature map is multiplied by its position attention weights to obtain the position attention feature map. In the channel dimension, Multiply the road feature map with the attention weights of each channel to obtain the channel attention feature map, and finally add the position attention feature map and the channel attention feature map to obtain the attention feature map corresponding to the road feature map.
例如,参阅图10所示,为本申请实施例中获得注意力特征图的示意图,假设道路特征图1的位置注意力权重矩阵为A,通道注意力权重矩阵为B,则将道路特征图1与位置注意力权重矩阵为A相乘,获得位置注意力特征图1,将道路特征图1与通道注意力权重矩阵为B相乘,获得通道注意力特征图1,最后将位置注意力特征图1和通道注意力特征图1相加,获得道路特征图1对应的注意力特征图1。For example, referring to Figure 10, it is a schematic diagram of obtaining the attention feature map in the embodiment of the present application. Assuming that the position attention weight matrix of the road feature map 1 is A, and the channel attention weight matrix is B, then the road feature map 1 Multiply with the position attention weight matrix A to obtain the position attention feature map 1, multiply the road feature map 1 with the channel attention weight matrix B to obtain the channel attention feature map 1, and finally the position attention feature map 1 is added to the channel attention feature map 1 to obtain the attention feature map 1 corresponding to the road feature map 1.
这样,基于一个道路特征图包含的各像素点各自的图像位置信息,结合一个道路特征图包含的图像通道信息,对道路特征图进行注意力加权处理,获得相应的注意力特征图,能够关注道路特征图中的重要信息,抑制无用信息,提高可行驶区域识别的准确度。In this way, based on the image position information of each pixel contained in a road feature map, combined with the image channel information contained in a road feature map, the attention weighting process is performed on the road feature map to obtain the corresponding attention feature map, which can pay attention to the road Important information in the feature map suppresses useless information and improves the accuracy of drivable area recognition.
步骤63:基于目标图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,并基于至少一个目标特征图,获得目标车辆的可行驶区域。Step 63: Based on the target image scale, perform image scale adjustment on the obtained attention feature maps to obtain at least one target feature map, and obtain the drivable area of the target vehicle based on the at least one target feature map.
本申请实施例中,在基于目标图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图时,服务器具体执行以下操作。参阅图11所示,其为本申请实施例获得至少一个目标特征图的流程示意图,下面结合附图11,对具体执行的操作进行详细说明:In the embodiment of the present application, when performing image scale adjustment on each of the obtained attention feature maps based on the target image scale to obtain at least one target feature map, the server specifically performs the following operations. Referring to FIG. 11 , it is a schematic flow diagram of obtaining at least one target feature map according to an embodiment of the present application. The specific operations performed will be described in detail below in conjunction with FIG. 11 :
步骤630:基于获得的各注意力特征图各自的图像尺度大小,对各注意力特征图进行融合更新,获得至少一个融合特征图。Step 630: Based on the obtained image scales of each attention feature map, perform fusion and update on each attention feature map to obtain at least one fusion feature map.
具体的,对各注意力特征图进行融合更新有多种方法。Specifically, there are many ways to fuse and update each attention feature map.
第一种方法:参阅图12所示,其为本申请实施例获得至少一个融合特征图的流程示意图,下面结合附图12,对具体执行的操作进行详细说明:The first method: refer to FIG. 12 , which is a schematic flow diagram of obtaining at least one fusion feature map according to the embodiment of the present application. The specific operations performed will be described in detail below in conjunction with FIG. 12 :
步骤6300:按照获得的各注意力特征图各自的图像尺度大小,将各注意力特征图进行排序,获得排序结果。Step 6300: According to the obtained image scales of each attention feature map, sort each attention feature map to obtain a sorting result.
本申请实施例中,按照获得的各注意力特征图各自的图像尺度大小,将各注意力特征图按照图像尺度从小到大的顺序进行排序,获得排序结果。In the embodiment of the present application, according to the respective image scales of the obtained attention feature maps, the attention feature maps are sorted in ascending order of image scale, and the sorting result is obtained.
例如,参阅图13所示,为本申请实施例中获得至少一个融合特征图的示意图,假设各注意力特征图包括注意力特征图P1、注意力特征图P2、注意力特征图P3、注意力特征图P4,注意力特征图P1的图像尺度大小为1/2、注意力特征图P2的图像尺度大小为1/4、注意力特征图P3的图像尺度大小为1/8、注意力特征图P4的图像尺度大小为1/16,则将各注意力特征图按照图像尺度从小到大进行排序,获得的排序结果为P1、P2、P3、P4。For example, referring to FIG. 13 , it is a schematic diagram of obtaining at least one fusion feature map in the embodiment of the present application. It is assumed that each attention feature map includes attention feature map P1, attention feature map P2, attention feature map P3, attention feature map Feature map P4, the image scale of the attention feature map P1 is 1/2, the image scale of the attention feature map P2 is 1/4, the image scale of the attention feature map P3 is 1/8, and the attention feature map The image scale of P4 is 1/16, and the attention feature maps are sorted according to the image scale from small to large, and the obtained sorting results are P1, P2, P3, and P4.
步骤6301:基于排序结果,依次读取每两个相邻的注意力特征图进行融合,直到读取完毕为止。Step 6301: Based on the sorting result, sequentially read every two adjacent attention feature maps for fusion until the reading is complete.
其中,一次融合过程包括:读取相邻的两个注意力特征图,并对两个注意力特征图进行融合,获得融合特征图。保存融合特征图,以及将融合特征图作为新的注意力特征图,代替两个注意力特征图,加入排序结果中。Wherein, one fusion process includes: reading two adjacent attention feature maps, and fusing the two attention feature maps to obtain the fusion feature map. Save the fusion feature map, and use the fusion feature map as a new attention feature map to replace the two attention feature maps and add it to the sorting result.
例如,如图13所示,样本排序结果为P1、P2、P3、P4。读取P1和P2,并对P1和P2进行融合,获得融合特征图R1,保存R1,以及将R1作为新的注意力特征图,代替P1和P2,加入排序结果中,此时排序结果为R1、P3、P4,读取R1和P3,并对R1和P3进行融合,获得融合特征图R2,保存R2,以及将R2作为新的注意力特征图,代替R1和P3,加入排序结果中,此时排序结果为R2、P4,读取R2和P4,并对R2和P4进行融合,获得融合特征图R3,保存R3,最终获得的至少一个融合特征图包括R1、R2和R3。For example, as shown in FIG. 13 , the sample sorting results are P1, P2, P3, and P4. Read P1 and P2, and fuse P1 and P2 to obtain the fusion feature map R1, save R1, and use R1 as a new attention feature map to replace P1 and P2, and add it to the sorting result, and the sorting result is now R1 , P3, P4, read R1 and P3, and fuse R1 and P3 to obtain the fusion feature map R2, save R2, and use R2 as a new attention feature map instead of R1 and P3, and add it to the sorting result. When the sorting results are R2 and P4, read R2 and P4, and fuse R2 and P4 to obtain the fusion feature map R3, save R3, and finally obtain at least one fusion feature map including R1, R2 and R3.
进一步地,不同图像尺度对应的注意力特征图之间的融合更新之前,将小图像尺度对应的道路特征图上采样至与大图像尺度对应的道路特征图相同的图像尺度,然后再基于上采样后的道路特征图包含的各像素点各自的图像位置信息,结合上采样后的道路特征图包含的图像通道信息,对上采样后的道路特征图进行注意力加权处理,获得小图像尺度对应的注意力特征图。Further, before the fusion update between the attention feature maps corresponding to different image scales, the road feature map corresponding to the small image scale is up-sampled to the same image scale as the road feature map corresponding to the large image scale, and then based on the up-sampling The image position information of each pixel contained in the post-sampled road feature map, combined with the image channel information contained in the up-sampled road feature map, performs attention weighting processing on the up-sampled road feature map, and obtains the corresponding small image scale. Attention feature map.
第二种方法:按照获得的各注意力特征图各自的图像尺度大小,将各注意力特征图进行排序,获得排序结果,并基于排序结果,依次读取每两个相邻的注意力特征图进行融合,直到读取完毕为止。The second method: according to the respective image scales of the obtained attention feature maps, sort the attention feature maps to obtain the sorting results, and based on the sorting results, read every two adjacent attention feature maps in turn Fusion is performed until the read is complete.
其中,一次融合过程包括:读取相邻的两个样本注意力特征图,并对该两个样本注意力特征图进行融合,获得样本融合特征图,保存样本融合特征图。Wherein, a fusion process includes: reading two adjacent sample attention feature maps, and fusing the two sample attention feature maps to obtain the sample fusion feature map, and saving the sample fusion feature map.
例如,假设各注意力特征图包括注意力特征图P1、注意力特征图P2、注意力特征图P3、注意力特征图P4,将各注意力特征图按照图像尺度从小到大进行排序,获得的排序结果为P1、P2、P3、P4,读取P1和P2,并对P1和P2进行融合,获得融合特征图R4,保存R4,读取P2和P3,并对P2和P3进行融合,获得融合特征图R5,保存R5,读取P3和P4,并对P3和P4进行融合,获得融合特征图R6,保存R6,最终获得的至少一个融合特征图包括R4、R5和R6。For example, assuming that each attention feature map includes attention feature map P1, attention feature map P2, attention feature map P3, and attention feature map P4, the attention feature maps are sorted according to the image scale from small to large, and the obtained The sorting results are P1, P2, P3, P4, read P1 and P2, and fuse P1 and P2 to obtain the fusion feature map R4, save R4, read P2 and P3, and fuse P2 and P3 to obtain the fusion Feature map R5, save R5, read P3 and P4, and fuse P3 and P4 to obtain a fusion feature map R6, save R6, and finally obtain at least one fusion feature map including R4, R5 and R6.
第三种方法:将各注意力特征图进行融合,获得一个融合特征图。The third method: fuse each attention feature map to obtain a fusion feature map.
例如,假设各注意力特征图包括注意力特征图P1、注意力特征图P2、注意力特征图P3、注意力特征图P4,则将P1、P2、P3、P4进行融合,获得一个融合特征图R7。For example, assuming that each attention feature map includes attention feature map P1, attention feature map P2, attention feature map P3, and attention feature map P4, then P1, P2, P3, and P4 are fused to obtain a fusion feature map R7.
步骤631:基于预设的目标图像通道数,分别对至少一个融合特征图进行图像通道调整,获得相应的中间特征图。Step 631: Based on the preset number of target image channels, perform image channel adjustment on at least one fused feature map to obtain corresponding intermediate feature maps.
本申请实施例中,获得至少一个融合特征图之后,针对至少一个融合特征图,分别执行以下操作:采用卷积层,基于预设的目标图像通道数,对一个融合特征图进行图像通道调整,获得该融合特征图对应的中间特征图。In the embodiment of the present application, after at least one fused feature map is obtained, the following operations are respectively performed for at least one fused feature map: using a convolutional layer, based on the preset number of target image channels, image channel adjustment is performed on a fused feature map, Obtain the intermediate feature map corresponding to the fused feature map.
其中,预设的目标图像通道数可以为目标道路图像的通道数,本申请实施例对此并不进行限制。Wherein, the preset target image channel number may be the channel number of the target road image, which is not limited in this embodiment of the present application.
例如,参阅图14所示,为本申请实施例中获得至少一个目标特征图的示意图,假设至少一个融合特征图包括融合特征图R1、融合特征图R2和融合特征图R3,预设的目标图像通道数为3,采用卷积层,分别对融合特征图R1、融合特征图R2和融合特征图R3进行图像通道调整,获得图像通道数为3的中间特征图Z1、图像通道数为3的中间特征图Z2、图像通道数为3的中间特征图Z3。For example, referring to Figure 14, which is a schematic diagram of obtaining at least one target feature map in the embodiment of the present application, assuming that at least one fusion feature map includes fusion feature map R1, fusion feature map R2 and fusion feature map R3, the preset target image The number of channels is 3, and the convolutional layer is used to adjust the image channels of the fusion feature map R1, fusion feature map R2 and fusion feature map R3 respectively, and obtain the intermediate feature map Z1 with 3 image channels and the intermediate feature map with 3 image channels. Feature map Z2, intermediate feature map Z3 with 3 image channels.
步骤632:基于目标图像尺度,对获得的至少一个中间特征图进行图像尺度调整,获得相应的目标特征图。Step 632: Based on the target image scale, perform image scale adjustment on at least one obtained intermediate feature map to obtain a corresponding target feature map.
本申请实施例中,获得至少一个中间特征图之后,针对至少一个中间特征图,分别执行以下操作:采用双线性插值,基于目标图像尺度,对一个中间特征图进行图像尺度调整,获得该中间特征图对应的目标特征图。In the embodiment of the present application, after at least one intermediate feature map is obtained, the following operations are performed on at least one intermediate feature map: using bilinear interpolation, based on the target image scale, an image scale is adjusted on an intermediate feature map to obtain the intermediate The target feature map corresponding to the feature map.
例如,如图14所示,至少一个中间特征图包括中间特征图Z1、中间特征图Z2和中间特征图Z3,目标图像尺度为640x480,采用双线性插值,分别对中间特征图Z1、中间特征图Z2和中间特征图Z3进行图像尺度调整,获得图像尺度为640x480的目标特征图M1、图像尺度为640x480的目标特征图M2、图像尺度为640x480的目标特征图M3。For example, as shown in Figure 14, at least one intermediate feature map includes intermediate feature map Z1, intermediate feature map Z2, and intermediate feature map Z3. Image scale adjustment is performed on image Z2 and intermediate feature map Z3 to obtain a target feature map M1 with an image scale of 640x480, a target feature map M2 with an image scale of 640x480, and a target feature map M3 with an image scale of 640x480.
这样,获得图像通道数和图像尺度都相同的至少一个样本目标特征图,整合了不同注意力特征图在通道和尺度特征上的差异,提高了可行驶区域识别的准确度。In this way, at least one sample target feature map with the same number of image channels and the same image scale is obtained, and the differences in channel and scale features of different attention feature maps are integrated to improve the accuracy of drivable area recognition.
本申请实施例中,在基于至少一个目标特征图,获得目标车辆的可行驶区域时,服务器具体执行以下操作。参阅图15所示,其为本申请实施例获得目标车辆的可行驶区域的流程示意图,下面结合附图15,对具体执行的操作进行详细说明:In the embodiment of the present application, when obtaining the drivable area of the target vehicle based on at least one target feature map, the server specifically performs the following operations. Referring to Figure 15, which is a schematic flow chart of obtaining the drivable area of the target vehicle in the embodiment of the present application, the specific operations performed will be described in detail below in conjunction with Figure 15:
步骤633:基于至少一个目标特征图,获得结果特征图。Step 633: Based on at least one target feature map, obtain a result feature map.
本申请实施例中,获得样本结果特征图,包含但不限于以下两种情况:In the embodiment of this application, the sample result feature map is obtained, including but not limited to the following two situations:
情况1:若至少一个目标特征图包括一个目标特征图,则将一个目标特征图作为结果特征图。Case 1: If at least one target feature map includes one target feature map, then use one target feature map as the result feature map.
例如,假设至少一个目标特征图只包括目标特征图M7,则将目标特征图M7作为结果特征图。For example, assuming that at least one target feature map only includes the target feature map M7, then the target feature map M7 is used as the result feature map.
情况2:若至少一个目标特征图包括多个目标特征图,则对多个目标特征图进行图像融合,获得结果特征图。Case 2: If at least one target feature map includes multiple target feature maps, image fusion is performed on the multiple target feature maps to obtain a resultant feature map.
例如,假设至少一个目标特征图包括目标特征图M1、目标特征图M2和目标特征图M3,则将M1、M2和M3进行融合,获得结果特征图。For example, assuming that at least one target feature map includes a target feature map M1, a target feature map M2, and a target feature map M3, M1, M2, and M3 are fused to obtain a resultant feature map.
步骤634:将结果特征图中包含的各像素点的值与预设的可行驶区域标签值进行比较,确定结果特征图中属于可行驶区域的各目标像素点。Step 634: Compare the value of each pixel contained in the result feature map with the preset drivable area label value, and determine each target pixel point belonging to the drivable area in the result feature map.
本申请实施例中,针对结果特征图中包含的各像素点,分别执行以下操作:将一个像素点的值与预设的可行驶区域标签值进行比较,若该像素点的值与可行驶区域标签值相同,则确定该像素点为目标像素点。In the embodiment of the present application, the following operations are performed for each pixel contained in the result feature map: compare the value of a pixel with the preset drivable area label value, If the label values are the same, the pixel is determined to be the target pixel.
其中,预设的可行驶区域标签值可以为1,本申请实施例对此并不进行限制。Wherein, the preset drivable area tag value may be 1, which is not limited in this embodiment of the present application.
例如,假设可行驶区域标签值为1,将结果特征图中包含的各像素点的值与标签值1进行比较,若结果特征图中与标签值1相同的各像素点分别为:像素点1、像素点2、像素点3,则确定结果特征图中属于可行驶区域的各目标像素点为像素点1、像素点2、像素点3。For example, assuming that the label value of the drivable area is 1, compare the value of each pixel contained in the result feature map with the label value 1, if the pixel points in the result feature map that are the same as the label value 1 are: pixel 1 , pixel 2, and pixel 3, then it is determined that each target pixel in the result feature map that belongs to the drivable area is pixel 1, pixel 2, and pixel 3.
步骤635:将目标道路图像中,与各目标像素点各自的图像位置对应的各像素点进行标示,获得包含可行驶区域的图像。Step 635: Mark each pixel point corresponding to the respective image position of each target pixel point in the target road image, and obtain an image including a drivable area.
例如,参阅图16所示,为本申请实施例中包含可行驶区域的图像的示意图,获得各目标像素点之后,将目标道路图像中,与各目标像素点各自的图像位置对应的各像素点进行标示,分割出可行驶区域,从而获得包含可行驶区域的图像。For example, referring to Figure 16, which is a schematic diagram of an image including a drivable area in the embodiment of the present application, after obtaining each target pixel point, each pixel point corresponding to the respective image position of each target pixel point in the target road image Mark and segment the drivable area to obtain an image containing the drivable area.
基于相同的发明构思,本申请实施例中还提供了一种可行驶区域识别装置,参阅图17所示,为本申请实施例中可行驶区域识别装置的结构示意图,具体包括:Based on the same inventive concept, the embodiment of the present application also provides a drivable area recognition device, as shown in Figure 17, which is a schematic structural diagram of the drivable area recognition device in the embodiment of the present application, specifically including:
获取模块1701,用于响应于驾驶对象触发的可行驶区域识别请求,获取当前待识别的目标道路图像,目标道路图像具有目标图像尺度;An acquisition module 1701, configured to acquire a current target road image to be identified in response to a drivable area identification request triggered by a driving object, where the target road image has a target image scale;
提取模块1702,用于按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图;The extraction module 1702 is used to perform feature extraction on the target road image according to the preset scales of each candidate image to obtain a corresponding road feature map;
第一处理模块1703,用于针对各道路特征图,分别执行以下操作:基于一个道路特征图包含的各像素点各自的图像位置信息,结合一个道路特征图包含的图像通道信息,对道路特征图进行注意力加权处理,获得相应的注意力特征图;The first processing module 1703 is configured to perform the following operations on each road feature map: based on the image position information of each pixel contained in a road feature map, combined with the image channel information contained in a road feature map, the road feature map Perform attention weighting processing to obtain the corresponding attention feature map;
第二处理模块1704,用于基于目标图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,并基于至少一个目标特征图,获得目标车辆的可行驶区域。The second processing module 1704 is configured to adjust the image scale of each attention feature map obtained based on the target image scale, obtain at least one target feature map, and obtain the drivable area of the target vehicle based on the at least one target feature map.
可选的,按照预设的各候选图像尺度,分别对目标道路图像进行特征提取,获得相应的道路特征图时,提取模块1702用于:Optionally, according to the preset scales of each candidate image, feature extraction is performed on the target road image respectively, and when the corresponding road feature map is obtained, the extraction module 1702 is used for:
针对预设的各候选图像尺度,分别执行以下操作:For each preset candidate image scale, perform the following operations:
基于一个候选图像尺度,对目标道路图像进行第一特征提取,获得第一特征图,第一特征图具有第一图像通道数;Based on a candidate image scale, performing first feature extraction on the target road image to obtain a first feature map, the first feature map has a first number of image channels;
按照预设的各候选图像感受野,分别对第一特征图进行第二特征提取,获得相应的第二特征图,并对获得的各第二特征图进行叠加,获得叠加特征图;其中,每个候选图像感受野表征:相应第二特征图包含的各像素点映射至目标道路图像上的区域大小;According to the preset receptive field of each candidate image, the second feature extraction is performed on the first feature map to obtain the corresponding second feature map, and the obtained second feature maps are superimposed to obtain the superimposed feature map; wherein, each A representation of the receptive field of a candidate image: each pixel contained in the corresponding second feature map is mapped to the area size on the target road image;
基于第一图像通道数,对叠加特征图进行图像通道调整,获得一个候选图像尺度对应的道路特征图。Based on the number of first image channels, image channel adjustment is performed on the superposition feature map to obtain a road feature map corresponding to a candidate image scale.
可选的,基于一个道路特征图包含的各像素点各自的图像位置信息,结合一个道路特征图包含的图像通道信息,对道路特征图进行注意力加权处理,获得相应的注意力特征图时,第一处理模块1703用于:Optionally, based on the respective image position information of each pixel contained in a road feature map, combined with the image channel information contained in a road feature map, the road feature map is subjected to attention weighting processing, and when the corresponding attention feature map is obtained, The first processing module 1703 is used for:
基于一个道路特征图包含的各像素点各自的图像位置信息,确定一个道路特征图包含的各像素点各自的位置注意力权重,其中,位置注意力权重表征在识别可行驶区域时,相应像素点的图像位置的重要程度;Based on the respective image position information of each pixel contained in a road feature map, determine the respective position attention weights of each pixel contained in a road feature map, where the position attention weight represents the corresponding pixel points when identifying the drivable area The importance of the image position;
基于一个道路特征图包含的图像通道信息,确定一个道路特征图包含的各图像通道各自的通道注意力权重,其中,通道注意力权重表征在识别可行驶区域时,相应图像通道的重要程度;Based on the image channel information contained in a road feature map, determine the respective channel attention weights of each image channel contained in a road feature map, wherein the channel attention weight represents the importance of the corresponding image channel when identifying the drivable area;
基于获得的各位置注意力权重和通道注意力权重,对一个道路特征图进行加权,获得一个道路特征图对应的注意力特征图。Based on the obtained attention weights of each position and channel attention weights, a road feature map is weighted to obtain an attention feature map corresponding to a road feature map.
可选的,基于目标图像尺度,对获得的各注意力特征图进行图像尺度调整,获得至少一个目标特征图,第二处理模块1704用于:Optionally, based on the target image scale, perform image scale adjustment on each of the obtained attention feature maps to obtain at least one target feature map, and the second processing module 1704 is used to:
基于获得的各注意力特征图各自的图像尺度大小,对各注意力特征图进行融合更新,获得至少一个融合特征图;Based on the respective image scales of the obtained attention feature maps, fusion and update are performed on each attention feature map to obtain at least one fusion feature map;
基于预设的目标图像通道数,分别对至少一个融合特征图进行图像通道调整,获得相应的中间特征图;Based on the preset number of target image channels, image channel adjustment is performed on at least one fusion feature map to obtain a corresponding intermediate feature map;
基于目标图像尺度,对获得的至少一个中间特征图进行图像尺度调整,获得相应的目标特征图。Based on the target image scale, image scale adjustment is performed on at least one obtained intermediate feature map to obtain a corresponding target feature map.
可选的,基于获得的各注意力特征图各自的图像尺度大小,对各注意力特征图进行融合更新,获得至少一个融合特征图时,第二处理模块1704还用于:Optionally, based on the respective image scales of the obtained attention feature maps, fusion and update are performed on each attention feature map, and when at least one fusion feature map is obtained, the second processing module 1704 is also used to:
按照获得的各注意力特征图各自的图像尺度大小,将各注意力特征图进行排序,获得排序结果;According to the respective image scales of the obtained attention feature maps, sort the attention feature maps to obtain the sorting results;
基于排序结果,依次读取每两个相邻的注意力特征图进行融合,直到读取完毕为止,其中,一次融合过程包括:Based on the sorting results, each two adjacent attention feature maps are read sequentially for fusion until the reading is completed. A fusion process includes:
读取相邻的两个注意力特征图,并对两个注意力特征图进行融合,获得融合特征图;Read two adjacent attention feature maps, and fuse the two attention feature maps to obtain a fusion feature map;
保存融合特征图,以及将融合特征图作为新的注意力特征图,代替两个注意力特征图,加入排序结果中。Save the fusion feature map, and use the fusion feature map as a new attention feature map to replace the two attention feature maps and add it to the sorting result.
可选的,基于至少一个目标特征图,获得目标车辆的可行驶区域时,第二处理模块1704还用于:Optionally, when obtaining the drivable area of the target vehicle based on at least one target feature map, the second processing module 1704 is also used to:
若至少一个目标特征图包括一个目标特征图,则将一个目标特征图作为结果特征图;If at least one target feature map includes a target feature map, then using a target feature map as a result feature map;
若至少一个目标特征图包括多个目标特征图,则对多个目标特征图进行图像融合,获得结果特征图;If at least one target feature map includes multiple target feature maps, performing image fusion on the multiple target feature maps to obtain a resultant feature map;
将结果特征图中包含的各像素点的值与预设的可行驶区域标签值进行比较,确定结果特征图中属于可行驶区域的各目标像素点;Comparing the value of each pixel contained in the result feature map with the preset drivable area label value, and determining each target pixel point belonging to the drivable area in the result feature map;
将目标道路图像中,与各目标像素点各自的图像位置对应的各像素点进行标示,获得包含可行驶区域的图像。In the target road image, each pixel point corresponding to the respective image position of each target pixel point is marked to obtain an image including a drivable area.
可选的,可行驶区域是通过将目标道路图像,输入目标识别模型获得的,装置还包括训练模块1705,训练模块1705用于:Optionally, the drivable area is obtained by inputting the target road image into the target recognition model, and the device further includes a training module 1705, which is used for:
基于训练样本集对待训练的识别模型进行迭代训练,获得目标识别模型;每个训练样本包括:样本道路的图像数据,其中,每个迭代过程执行以下操作:Perform iterative training on the recognition model to be trained based on the training sample set to obtain the target recognition model; each training sample includes: image data of the sample road, wherein each iteration process performs the following operations:
按照预设的各样本图像尺度,分别对选取的训练样本进行特征提取,获得相应的样本特征图,训练样本具有样本图像尺度;According to the preset image scales of each sample, feature extraction is performed on the selected training samples to obtain corresponding sample feature maps, and the training samples have sample image scales;
针对各样本特征图,分别执行以下操作:基于一个样本特征图包含的各像素点各自的样本图像位置信息,结合一个样本特征图包含的样本图像通道信息,对样本特征图进行注意力加权处理,获得相应的样本注意力特征图;For each sample feature map, perform the following operations respectively: Based on the respective sample image position information of each pixel contained in a sample feature map, combined with the sample image channel information contained in a sample feature map, perform attention weighting processing on the sample feature map, Obtain the corresponding sample attention feature map;
基于样本图像尺度,分别对获得的各样本注意力特征图进行图像尺度调整,获得至少一个目标样本特征图;Based on the sample image scale, image scale adjustment is performed on the obtained attention feature maps of each sample to obtain at least one target sample feature map;
基于至少一个目标样本特征图,获得训练样本的可行驶区域预测结果,并基于可行驶区域预测结果对应的损失值进行调参。Based on at least one target sample feature map, the drivable area prediction result of the training sample is obtained, and the parameters are adjusted based on the loss value corresponding to the drivable area prediction result.
基于上述实施例,参阅图18所示为本申请实施例中电子设备的结构示意图。Based on the above embodiments, refer to FIG. 18 , which is a schematic structural diagram of an electronic device in an embodiment of the present application.
本申请实施例提供了一种电子设备,该电子设备可以包括处理器1810(CenterProcessing Unit,CPU)、存储器1820、输入设备1830和输出设备1840等,输入设备1830可以包括键盘、鼠标、触摸屏等,输出设备1840可以包括显示设备,如液晶显示器(LiquidCrystal Display,LCD)、阴极射线管(Cathode Ray Tube,CRT)等。An embodiment of the present application provides an electronic device, which may include a processor 1810 (Center Processing Unit, CPU), a memory 1820, an input device 1830, an output device 1840, etc., and the input device 1830 may include a keyboard, a mouse, a touch screen, etc., The output device 1840 may include a display device, such as a liquid crystal display (Liquid Crystal Display, LCD), a cathode ray tube (Cathode Ray Tube, CRT), and the like.
存储器1820可以包括只读存储器(ROM)和随机存取存储器(RAM),并向处理器1810提供存储器1820中存储的程序指令和数据。在本申请实施例中,存储器1820可以用于存储本申请实施例中任一种可行驶区域识别方法的程序。The memory 1820 may include Read Only Memory (ROM) and Random Access Memory (RAM), and provides program instructions and data stored in the memory 1820 to the processor 1810 . In the embodiment of the present application, the memory 1820 may be used to store programs of any drivable area identification method in the embodiment of the present application.
处理器1810通过调用存储器1820存储的程序指令,处理器1810用于按照获得的程序指令执行本申请实施例中任一种可行驶区域识别方法。The processor 1810 invokes the program instructions stored in the memory 1820, and the processor 1810 is configured to execute any drivable area identification method in the embodiments of the present application according to the obtained program instructions.
基于上述实施例,本申请实施例中,提供了一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现上述任意方法实施例中的可行驶区域识别方法。Based on the above-mentioned embodiments, in the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the drivable area identification method in any of the above-mentioned method embodiments is implemented .
本领域内的技术人员应明白,本申请的实施例可提供为方法、系统、或计算机程序产品。因此,本申请可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本申请可采用在一个或多个其中包含有计算机可用程序代码的计算机可用存储介质(包括但不限于磁盘存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。Those skilled in the art should understand that the embodiments of the present application may be provided as methods, systems, or computer program products. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application may take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) having computer-usable program code embodied therein.
本申请是参照根据本申请的方法、设备(系统)、和计算机程序产品的流程图和/或方框图来描述的。应理解可由计算机程序指令实现流程图和/或方框图中的每一流程和/或方框、以及流程图和/或方框图中的流程和/或方框的结合。可提供这些计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理设备的处理器以产生一个机器,使得通过计算机或其他可编程数据处理设备的处理器执行的指令产生用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的装置。The present application is described with reference to flowchart illustrations and/or block diagrams of methods, apparatus (systems), and computer program products according to the present application. It should be understood that each procedure and/or block in the flowchart and/or block diagram, and a combination of procedures and/or blocks in the flowchart and/or block diagram can be realized by computer program instructions. These computer program instructions may be provided to a general purpose computer, special purpose computer, embedded processor, or processor of other programmable data processing equipment to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing equipment produce a An apparatus for realizing the functions specified in one or more procedures of the flowchart and/or one or more blocks of the block diagram.
这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理设备以特定方式工作的计算机可读存储器中,使得存储在该计算机可读存储器中的指令产生包括指令装置的制造品,该指令装置实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能。These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing apparatus to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture comprising instruction means, the instructions The device realizes the function specified in one or more procedures of the flowchart and/or one or more blocks of the block diagram.
这些计算机程序指令也可装载到计算机或其他可编程数据处理设备上,使得在计算机或其他可编程设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程设备上执行的指令提供用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的步骤。These computer program instructions can also be loaded onto a computer or other programmable data processing device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby The instructions provide steps for implementing the functions specified in the flow chart or blocks of the flowchart and/or the block or blocks of the block diagrams.
显然,本领域的技术人员可以对本申请进行各种改动和变型而不脱离本申请的精神和范围。这样,倘若本申请的这些修改和变型属于本申请权利要求及其等同技术的范围之内,则本申请也意图包含这些改动和变型在内。Obviously, those skilled in the art can make various changes and modifications to the application without departing from the spirit and scope of the application. In this way, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims (10)
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310504926.3A CN116543369A (en) | 2023-05-04 | 2023-05-04 | Method and device for identifying drivable area, electronic equipment and storage medium |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310504926.3A CN116543369A (en) | 2023-05-04 | 2023-05-04 | Method and device for identifying drivable area, electronic equipment and storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| CN116543369A true CN116543369A (en) | 2023-08-04 |
Family
ID=87443004
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CN202310504926.3A Pending CN116543369A (en) | 2023-05-04 | 2023-05-04 | Method and device for identifying drivable area, electronic equipment and storage medium |
Country Status (1)
| Country | Link |
|---|---|
| CN (1) | CN116543369A (en) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200064855A1 (en) * | 2018-08-27 | 2020-02-27 | Samsung Electronics Co., Ltd. | Method and apparatus for determining road line |
| CN114674338A (en) * | 2022-04-08 | 2022-06-28 | 石家庄铁道大学 | Refined recommendation method of road drivable area based on hierarchical input-output and dual-attention jumping |
| CN114694115A (en) * | 2022-03-24 | 2022-07-01 | 商汤集团有限公司 | Road obstacle detection method, device, equipment and storage medium |
| CN115620017A (en) * | 2022-09-30 | 2023-01-17 | 中汽创智科技有限公司 | Image feature extraction method, device, equipment and storage medium |
| US20230069215A1 (en) * | 2021-08-27 | 2023-03-02 | Motional Ad Llc | Navigation with Drivable Area Detection |
-
2023
- 2023-05-04 CN CN202310504926.3A patent/CN116543369A/en active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200064855A1 (en) * | 2018-08-27 | 2020-02-27 | Samsung Electronics Co., Ltd. | Method and apparatus for determining road line |
| US20230069215A1 (en) * | 2021-08-27 | 2023-03-02 | Motional Ad Llc | Navigation with Drivable Area Detection |
| CN114694115A (en) * | 2022-03-24 | 2022-07-01 | 商汤集团有限公司 | Road obstacle detection method, device, equipment and storage medium |
| CN114674338A (en) * | 2022-04-08 | 2022-06-28 | 石家庄铁道大学 | Refined recommendation method of road drivable area based on hierarchical input-output and dual-attention jumping |
| CN115620017A (en) * | 2022-09-30 | 2023-01-17 | 中汽创智科技有限公司 | Image feature extraction method, device, equipment and storage medium |
Non-Patent Citations (2)
| Title |
|---|
| ZENGYU QIU等: "MFIALane: Multiscale Feature Information Aggregator Network for Lane Detection", 《IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS》, vol. 23, no. 12, 31 August 2022 (2022-08-31), pages 24263 - 24275 * |
| 黄篷迟: "无人驾驶视觉环境感知目标检测与分割技术研究", 《 中国优秀硕士学位论文全文数据库 (工程科技Ⅱ辑)》, no. 2023, 15 January 2023 (2023-01-15), pages 035 - 1177 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN113139543B (en) | Training method of target object detection model, target object detection method and device | |
| JP2023527615A (en) | Target object detection model training method, target object detection method, device, electronic device, storage medium and computer program | |
| CN111126258A (en) | Image recognition method and related device | |
| CN111860398A (en) | Remote sensing image target detection method, system and terminal device | |
| CN114385662B (en) | Road network updating method and device, storage medium and electronic equipment | |
| CN114419338B (en) | Image processing method, image processing device, computer equipment and storage medium | |
| CN116645592B (en) | Crack detection method based on image processing and storage medium | |
| CN112580558A (en) | Infrared image target detection model construction method, detection method, device and system | |
| CN113837155B (en) | Image processing method, map data updating device and storage medium | |
| CN111062964A (en) | Image segmentation method and related device | |
| CN115131282B (en) | Target detection method, device, storage medium, equipment and program product | |
| CN116012626A (en) | Material matching method, device, equipment and storage medium for building facade images | |
| CN119152185A (en) | Target detection method, device, vehicle and storage medium | |
| CN114359231A (en) | Parking space detection method, device, equipment and storage medium | |
| CN116628531B (en) | Crowd-sourced map road object element clustering method, system and storage medium | |
| CN119731689A (en) | Image processing method, model training method, device and terminal equipment | |
| CN118864829B (en) | Object detection method and device, storage medium and electronic equipment | |
| CN118246511B (en) | A training method, system, device and medium for vehicle detection model | |
| CN120388334A (en) | Road image detection method, device and storage medium | |
| CN116543369A (en) | Method and device for identifying drivable area, electronic equipment and storage medium | |
| CN112434591A (en) | Lane line determination method and device | |
| CN115424027B (en) | Image similarity comparison method, device and equipment for image foreground person | |
| CN116229130A (en) | Fuzzy image type identification method, device, computer equipment and storage medium | |
| CN114596698B (en) | Road monitoring equipment position judging method, device, storage medium and equipment | |
| CN116092031A (en) | Method and device for extracting dotted line lane line based on optical remote sensing image |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PB01 | Publication | ||
| PB01 | Publication | ||
| SE01 | Entry into force of request for substantive examination | ||
| SE01 | Entry into force of request for substantive examination |