Detailed Description
The present invention will be described in further detail with reference to the drawings and examples, in order to make the objects, technical solutions and advantages of the present invention more apparent. It should be understood that the specific embodiments described herein are for purposes of illustration only and are not intended to limit the scope of the invention.
The foggy day image data set synthesis method based on depth estimation provided by the invention can be applied to a foggy image synthesis frame shown in fig. 1, and comprises three modules, namely a depth estimation model, an optical foggy model and a synthetic foggy data set, and the depth estimation of an image scene, the foggy processing of the optical model and the pairing and arrangement of the synthetic foggy data are respectively realized.
In one embodiment, as shown in fig. 2, there is provided a foggy day image data set synthesizing method based on depth estimation, which is described by taking an example that the method is applied to the foggy image synthesizing frame in fig. 1, and includes the following steps:
and 202, constructing a foggy day image data synthesis model.
The foggy day image data synthesis model comprises a depth estimation model, an optical foggy adding model and a synthetic foggy image model.
And 204, after the depth estimation model is optimally trained by adopting a monocular depth estimation method according to the public data set, inputting the target image scene data into the optimized depth estimation model for semantic deepening, and outputting the synthetic fog effect of the depth map corresponding to the target image scene data.
And 206, generating an atomized image after performing atomization treatment on the depth map through the optical atomization model according to the synthetic fog effect.
And step 208, determining a foggy visual task matching algorithm of the synthetic foggy image model according to the foggy image and a preset target task, and generating a target foggy image data set according to the foggy visual task matching algorithm.
In the foggy day image data set synthesis method based on depth estimation, the optical phenomenon in the foggy day environment can be more accurately simulated by introducing the combination mode of the depth estimation model and the optical foggy model. After the depth estimation model is optimized and trained, the depth information of the target image is obtained by using the model, and semantic deepening is carried out, so that a depth map which is more in line with the real world fog effect is generated. Then, the target image is atomized according to the depth information of the scene through an optical fogging model, and an image with a natural fogging effect is generated. And secondly, the atomization process of the optical atomization model is guided by the depth map information, so that the atomization effect which is finer and adapts to the complex outdoor environment is realized. Specifically, the depth map can provide spatial position information of different objects and backgrounds in the scene, so that the optical fog adding model generates fog layers with different densities and layers according to the actual structure of the scene, and a more real fog effect is simulated. In addition, the technical scheme also solves the problems of the virtual scene simulation fog technology in the aspects of generating the authenticity and the universality of the image data. Through the combined use of the depth estimation and the optical fogging model, the foggy image can be synthesized in the real scene, and the limitation of the virtual scene is avoided. In addition, the synthetic fog image model also presets a target task to determine a fog visual task matching algorithm according to the characteristics of the original public data set, so that a target fog image data set with the original detailed label is generated, the quality and the application universality of the data set are further improved, and the fog image data set with high quality and the detailed label is successfully constructed. The data set not only can improve the application effect of the image processing and the computer vision model in the foggy environment, but also can provide more real and efficient data support for tasks such as target detection and fusion in the foggy scene.
In one embodiment, the depth estimation model is built based on ResNet, wherein the depth estimation model includes a feature fusion block, a residual convolution block, and an adaptive output block.
It is worth noting that a depth estimation model is used to generate a depth map of the target image scene. The depth estimation model is shown in fig. 3. The model is built based on ResNet frames, is formed by combining a feature fusion plate, a residual convolution plate and a self-adaptive output plate, and can realize conversion from an input image to a corresponding depth map. The model is trained using large-scale data of multiple classical datasets using a monocular depth estimation method. In particular, training of each dataset is considered an independent task aimed at seeking an approximate pareto optimal solution on the dataset.
In one embodiment, a monocular depth estimation method is used to define a loss function for each of the training data sets trained on the depth estimation model based on different training data sets in the public data set:
;
;
wherein, the For the size of the training data set i,For scale and translational invariance loss,For a gradient matching item with multiple scales and unchanged scales, M is the number of pixel points contained in the corresponding image of the training data set,As the disparity prediction value, a disparity value,As the true value of the tag is,As a type of loss function,Is a super parameter. And minimizing the loss function by adopting a least square method to obtain a scale and translation invariant loss function:
;
;
where s is a resolution parameter, t is a position parameter, To normalize the disparity prediction value to the label value difference,For the disparity map differences at scale k,To be aligned in the x-axis directionIs used for the gradient of (a),To be aligned in the y-axis directionIs used for the gradient of (a),Is a multi-scale varying scale level.
And enhancing the training task by introducing a 3D data set into each training data set, and taking the minimum value of all the trained scales and translation invariant loss functions as a model optimization parameter if the scale and translation invariant loss functions of the optimized training according to the enhanced task are reduced:
;
wherein, the For the enhanced task of the current training,Is a model parameter. Otherwise, the scale and translation invariable loss function is reserved, and the current trained enhanced task is discarded, so that a task data set to be optimized is obtained. And optimizing the depth estimation model according to the scale and translation invariant loss function and the task data set to be optimized, and training a target image through the optimized depth estimation model.
In one embodiment, inputting target image scene data and the target image into an optimized depth estimation model for semantic deepening, outputting a predicted depth map, and performing brightness inversion on the predicted depth map after gray level conversion to generate a first gray level map;
Extracting scene depth information of the first gray scale map by adopting a standard optical algorithm:
;
;
wherein, the In order to atomize the image,In order to be an image of the object,In the case of a transmission diagram,Is the light of the atmosphere, and is the light of the atmosphere,For a pixel point in the image,For the depth of the scene,Is the atmospheric scattering coefficient. And generating a synthetic fog effect of the depth map corresponding to the target image scene data according to the scene depth information.
It should be noted that, based on the standard optical model, the atomized image is synthesized by the depth map and the original clear image, and the specific process is shown in fig. 4. Firstly, gray level conversion and brightness inversion are carried out on the depth map obtained in the previous stage, and a corresponding gray level map is generated. Subsequently, the optical formula is used:
;
wherein, the In order to atomize the image,In order to be an image of the object,In the case of a transmission diagram,Is the light of the atmosphere, and is the light of the atmosphere,Is a pixel point in the image. The transmission map can be represented by the formulaWhereinThe depth of the scene is represented as,Is the atmospheric scattering coefficient.
In one embodiment, after the visible light image of the depth map is subjected to superposition atomization processing through the optical atomization model according to the synthetic fog effect, an atomization image corresponding to the depth map is generated.
In one embodiment, the target tasks include constructing a foggy visual task matching algorithm performance test dataset, constructing a foggy infrared-visible pixel level pairing fusion dataset, and constructing a foggy target detection dataset.
In one embodiment, a performance test data set of a foggy visual task matching algorithm is constructed, and the method specifically comprises the steps of selecting MFNet semantic segmentation parts of a public data set, and firstly extracting and reconstructing 4-channel data of the semantic segmentation parts to obtain a 3-channel RGB image. Then, the image is classified according to the original classification label into daytime and night scenes. The classified data are input into a depth estimation model, corresponding fog brightness is matched, and fog image data sets with different concentrations in daytime and night are generated by adjusting fog concentration parameters. The data set reserves the original image information and is matched with the depth map, so that the data set is suitable for effect test and contrast evaluation of a foggy day visual task matching algorithm, and subjective visual comparison and objective index evaluation are supported. Examples of partial datasets are shown in fig. 5 and 6.
It is worth to say that, by adopting the method, the synthetic foggy day images with multiple scenes and different concentrations and the corresponding depth map can be quickly generated only by inputting clear images.
In one embodiment, a foggy infrared-visible pixel level paired fusion dataset is constructed by conducting an experiment based on RoadScene public datasets. The dataset independently divides the infrared and visible images. Firstly classifying 221 visible light images, dividing the visible light images into a daytime scene under a strong illumination condition and a night scene under a weak illumination condition, and generating a corresponding depth map by using a depth estimation model. In view of the fact that the daytime scene in the data set is sufficient in sunlight and the night scene is more in light, a higher atomization brightness value is selected to be matched with scene conditions. Finally, the atomized visible light image is paired with the original infrared image, and a data set comprising the visible light image, the infrared image, the depth map and the atomized image which are paired at the pixel level is constructed, so that the method is suitable for an image fusion task. At the same time, preserving the original image facilitates evaluation of fusion algorithm based on subjective vision and objective indexes, such as contrast analysis of indexes of Visual Information Fidelity (VIF) and mutual information quantity (MI). The foggy-day infrared-visible light pixel level pairing fusion dataset constructed based on RoadScene dataset shows that the dataset realizes pairing of pixel levels, namely a visible light image, a depth map, an atomized image and an infrared image from left to right as shown in fig. 7 and 8.
It is worth to say that the synthetic fog superposition is performed on the visible light image based on the existing visible light-infrared paired data set. The investigation shows that the foggy weather has no influence on infrared imaging basically, so that the original infrared image data and labels are reserved, and a visible light-infrared pixel level pairing data set under foggy conditions is constructed. The data set can effectively fill the blank of the foggy-day multi-mode study, and provides a large amount of reliable test data for the related image fusion algorithm.
In one embodiment, a foggy-day target detection dataset is constructed by selecting MFNet detection parts in the public dataset, and classifying the detection parts according to illumination conditions, wherein the detection parts are classified into strong illumination, insufficient illumination and weak illumination as shown in fig. 9. The visible light image data in the categories are input into a depth estimation model, after the corresponding fog brightness is matched, different fog concentration parameters are selected to generate fog day visible light image data with different fog concentrations under the illumination condition shown in fig. 10, and after the fog day visible light image data is paired with an infrared image in the original data set and a detection label, a fog day target detection data set is obtained. Fig. 10 is a partial image data comparison display of the dataset, from top to bottom, of a day fog, a mist scene, a night fog, and a mist scene.
It is worth to say that the fog adding processing is carried out on the existing marked public data set, and a foggy day image with clear marks is generated so as to support training and generalization of a foggy day target detection algorithm. Meanwhile, as the interference of the infrared image on fog is small, the infrared and visible light paired data set is selected and utilized to evaluate the performance of the detection algorithm on different image data, so that the overall accuracy of the detection algorithm is improved.
It should be understood that, although the steps in the flowcharts of fig. 2 and 4 are shown in order as indicated by the arrows, these steps are not necessarily performed in order as indicated by the arrows. The steps are not strictly limited to the order of execution unless explicitly recited herein, and the steps may be executed in other orders. Moreover, at least some of the steps in fig. 2, 4 may comprise a plurality of sub-steps or phases, which are not necessarily performed at the same time, but may be performed at different times, which are not necessarily performed sequentially, but may be performed alternately or alternately with at least some of the other steps or sub-steps of other steps.
In one embodiment, as shown in FIG. 11, there is provided a foggy day image data integration apparatus based on depth estimation, comprising a synthetic model construction module 1102, a synthetic fog effect generation module 1104, a foggy image generation module 1106, and a synthetic fog data module 1108, wherein:
The synthetic model construction module 1102 is configured to construct a foggy day image data synthetic model. The foggy day image data synthesis model comprises a depth estimation model, an optical foggy adding model and a synthetic foggy image model.
The synthetic fog effect generating module 1104 is configured to perform optimization training on the depth estimation model by using a monocular depth estimation method according to the public data set, input the target image scene data to the optimized depth estimation model for semantic deepening, and output the synthetic fog effect of the depth map corresponding to the target image scene data.
And the atomized image generating module 1106 is configured to generate an atomized image after performing atomization processing on the depth map through the optical atomization model according to the synthetic fog effect.
The synthetic fog data module 1108 is configured to determine a foggy day visual task matching algorithm of the synthetic fog image model according to the foggy image and a preset target task, and generate a target foggy day image dataset according to the foggy day visual task matching algorithm.
For a detailed definition of a foggy day image data set synthesizing device based on depth estimation, reference may be made to the definition of a foggy day image data set synthesizing method based on depth estimation hereinabove, and the detailed description thereof will be omitted. The foregoing module of the foggy day image data integration device based on depth estimation may be implemented in whole or in part by software, hardware, or a combination thereof. The above modules may be embedded in hardware or may be independent of a processor in the computer device, or may be stored in software in a memory in the computer device, so that the processor may call and execute operations corresponding to the above modules.
In one embodiment, a computer device is provided, which may be a terminal, and the internal structure thereof may be as shown in fig. 12. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Wherein the processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface of the computer device is used for communicating with an external terminal through a network connection. The computer program, when executed by a processor, implements a foggy day image data set synthesis method based on depth estimation. The display screen of the computer equipment can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer equipment can be a touch layer covered on the display screen, can also be keys, a track ball or a touch pad arranged on the shell of the computer equipment, and can also be an external keyboard, a touch pad or a mouse and the like.
It will be appreciated by persons skilled in the art that the structures shown in fig. 11-12 are block diagrams of only portions of structures associated with the present inventive arrangements and are not limiting of the computer device to which the present inventive arrangements may be implemented, and that a particular computer device may include more or fewer components than shown, or may be combined with certain components, or have different arrangements of components.
In one embodiment, a computer device is provided comprising a memory storing a computer program and a processor that when executing the computer program performs the steps of:
and constructing a foggy day image data synthesis model. The foggy day image data synthesis model comprises a depth estimation model, an optical foggy adding model and a synthetic foggy image model.
And (3) after the depth estimation model is optimized and trained by adopting a monocular depth estimation method according to the public data set, inputting the target image scene data into the optimized depth estimation model for semantic deepening, and outputting the synthetic fog effect of the depth map corresponding to the target image scene data.
And generating an atomized image after performing atomization treatment on the depth map through the optical atomization model according to the synthetic fog effect.
Determining a foggy visual task matching algorithm for synthesizing a foggy image model according to the foggy image and a preset target task, and generating a target foggy image data set according to the foggy visual task matching algorithm
Those skilled in the art will appreciate that implementing all or part of the above described methods may be accomplished by way of a computer program stored on a non-transitory computer readable storage medium, which when executed, may comprise the steps of the embodiments of the methods described above. Any reference to memory, storage, database, or other medium used in embodiments provided herein may include non-volatile and/or volatile memory. The nonvolatile memory can include Read Only Memory (ROM), programmable ROM (PROM), electrically Programmable ROM (EPROM), electrically Erasable Programmable ROM (EEPROM), or flash memory. Volatile memory can include Random Access Memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms such as Static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double Data Rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (SYNCHLINK) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), among others.
The technical features of the above embodiments may be arbitrarily combined, and all possible combinations of the technical features in the above embodiments are not described for brevity of description, however, as long as there is no contradiction between the combinations of the technical features, they should be considered as the scope of the description.
The above examples illustrate only a few embodiments of the invention, which are described in detail and are not to be construed as limiting the scope of the invention. It should be noted that it will be apparent to those skilled in the art that several variations and modifications can be made without departing from the spirit of the invention, which are all within the scope of the invention. Accordingly, the scope of protection of the present invention is to be determined by the appended claims.