Deprecated: The each() function is deprecated. This message will be suppressed on further calls in /home/zhenxiangba/zhenxiangba.com/public_html/phproxy-improved-master/index.php on line 456
CN120451727B - A method and device for synthesizing foggy image data sets based on depth estimation - Google Patents
[go: Go Back, main page]

CN120451727B - A method and device for synthesizing foggy image data sets based on depth estimation - Google Patents

A method and device for synthesizing foggy image data sets based on depth estimation

Info

Publication number
CN120451727B
CN120451727B CN202510948609.XA CN202510948609A CN120451727B CN 120451727 B CN120451727 B CN 120451727B CN 202510948609 A CN202510948609 A CN 202510948609A CN 120451727 B CN120451727 B CN 120451727B
Authority
CN
China
Prior art keywords
image
model
foggy
depth estimation
depth
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Active
Application number
CN202510948609.XA
Other languages
Chinese (zh)
Other versions
CN120451727A (en
Inventor
罗澜
张志远
师俊朋
黄可馨
郑雯
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
National University of Defense Technology
Original Assignee
National University of Defense Technology
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by National University of Defense Technology filed Critical National University of Defense Technology
Priority to CN202510948609.XA priority Critical patent/CN120451727B/en
Publication of CN120451727A publication Critical patent/CN120451727A/en
Application granted granted Critical
Publication of CN120451727B publication Critical patent/CN120451727B/en
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02ATECHNOLOGIES FOR ADAPTATION TO CLIMATE CHANGE
    • Y02A90/00Technologies having an indirect contribution to adaptation to climate change
    • Y02A90/10Information and communication technologies [ICT] supporting adaptation to climate change, e.g. for weather forecasting or climate simulation

Landscapes

  • Image Processing (AREA)

Abstract

本发明涉及一种基于深度估计的雾天图像数据集合成方法及装置。所述方法包括:构建雾天图像数据合成模型。其包括:深度估计模型、光学加雾模型以及合成雾图像模型。根据公开数据集采用单目深度估计方法对深度估计模型进行优化训练后,将目标图像场景数据输入至已优化的深度估计模型进行语义深化,输出目标图像场景数据对应深度图的合成雾效应。根据合成雾效应通过光学加雾模型对深度图进行雾化处理后,生成雾化图像。根据雾化图像与预设的目标任务确定合成雾图像模型的雾天视觉任务匹配算法,根据雾天视觉任务匹配算法生成目标雾天图像数据集。采用本方法能够快速构建大规模、多种光照条件和不同浓度的合成雾数据集,以支持高级视觉处理任务。

The present invention relates to a method and device for synthesizing a foggy image dataset based on depth estimation. The method comprises: constructing a foggy image data synthesis model. The method comprises: a depth estimation model, an optical fogging model, and a synthetic fog image model. After optimizing and training the depth estimation model using a monocular depth estimation method according to a public dataset, the target image scene data is input into the optimized depth estimation model for semantic deepening, and a synthetic fog effect of a depth map corresponding to the target image scene data is output. After the depth map is atomized using an optical fogging model according to the synthetic fog effect, a foggy image is generated. The foggy visual task matching algorithm of the synthetic fog image model is determined based on the foggy image and a preset target task, and the target foggy image dataset is generated based on the foggy visual task matching algorithm. The present method can quickly construct a large-scale synthetic fog dataset with multiple lighting conditions and different concentrations to support advanced visual processing tasks.

Description

Foggy day image data set synthesis method and device based on depth estimation
Technical Field
The invention relates to the technical field of fog data integration, in particular to a method and a device for foggy day image data integration based on depth estimation.
Background
Artificial intelligence is increasingly integrated into our lives, and industries such as automatic driving, monitoring and detection are required to realize comprehensive development, and have adaptability under various weather conditions. Fog is a common natural phenomenon, particularly in the early morning of autumn and winter and in the mountains in the south, but the existing foggy day data set is limited in quantity and uneven in data quality due to illumination and environmental limitations. The concentration of the mist and the lighting conditions are not uniform, resulting in poor usability. After the multi-mode fusion technology is introduced, the infrared and visible light images are paid attention to due to complementarity, and because the paired data sets in foggy days are scarce and have great real acquisition difficulty, the synthesized fog data set has remarkable advantages, and can better support advanced visual task models such as fusion, detection and the like.
In the acquisition of the foggy data set, the acquisition of real data is difficult and the updating cost of the tag is high. In order to solve the problem of scarcity of open source data, a foggy day image data set is often constructed by superposing virtual fogs on an image. However, the prior art method has a plurality of defects in the effect of generating fog, namely firstly, the existing RGB channel-based synthetic fog technology only applies a uniform density fog layer on an image, lacks modeling application to natural optical laws, and has unrealistic generating effect, secondly, the center point synthetic fog method based on a standard optical model is applied to an atmospheric scattering model, but has a limited usable range due to random selection of a center point, and is difficult to be applied to a complex outdoor environment, and in addition, the virtual scene simulation fog technology can generate a paired image on a simulation platform, but the authenticity and universality of generated image data still need to be verified, and meanwhile, the application of labels in fusion and detection models is limited due to the defects, so that the quality and the practical application effect of a synthetic fog data set are affected.
Disclosure of Invention
In view of the foregoing, it is desirable to provide a method and apparatus for synthesizing a foggy image dataset based on depth estimation, which can improve the accuracy of generating a synthetic foggy dataset and the efficiency of constructing a large-scale dataset.
A method of foggy day image dataset synthesis based on depth estimation, the method comprising:
and constructing a foggy day image data synthesis model. The foggy day image data synthesis model comprises a depth estimation model, an optical foggy adding model and a synthetic foggy image model.
And (3) after the depth estimation model is optimized and trained by adopting a monocular depth estimation method according to the public data set, inputting the target image scene data into the optimized depth estimation model for semantic deepening, and outputting the synthetic fog effect of the depth map corresponding to the target image scene data.
And generating an atomized image after performing atomization treatment on the depth map through the optical atomization model according to the synthetic fog effect.
And determining a foggy visual task matching algorithm for synthesizing a foggy image model according to the foggy image and a preset target task, and generating a target foggy image data set according to the foggy visual task matching algorithm.
A foggy day image data aggregation apparatus based on depth estimation, the apparatus comprising:
and the synthesis model construction module is used for constructing a foggy day image data synthesis model. The foggy day image data synthesis model comprises a depth estimation model, an optical foggy adding model and a synthetic foggy image model.
And the synthetic fog effect generation module is used for inputting the target image scene data into the optimized depth estimation model for semantic deepening after optimizing and training the depth estimation model by adopting a monocular depth estimation method according to the public data set, and outputting the synthetic fog effect of the depth map corresponding to the target image scene data.
And the atomized image generation module is used for generating an atomized image after performing atomization treatment on the depth map through the optical atomization model according to the synthetic fog effect.
And the synthetic fog data module is used for determining a foggy day visual task matching algorithm of the synthetic fog image model according to the foggy image and a preset target task, and generating a target foggy day image data set according to the foggy day visual task matching algorithm.
According to the foggy day image data set synthesis method and device based on depth estimation, the optical phenomenon in the foggy day environment can be more accurately simulated by introducing the combination mode of the depth estimation model and the optical foggy model. After the depth estimation model is optimized and trained, the depth information of the target image is obtained by using the model, and semantic deepening is carried out, so that a depth map which is more in line with the real world fog effect is generated. Then, the target image is atomized according to the depth information of the scene through an optical fogging model, and an image with a natural fogging effect is generated. And secondly, the atomization process of the optical atomization model is guided by the depth map information, so that the atomization effect which is finer and adapts to the complex outdoor environment is realized. Specifically, the depth map can provide spatial position information of different objects and backgrounds in the scene, so that the optical fog adding model generates fog layers with different densities and layers according to the actual structure of the scene, and a more real fog effect is simulated. In addition, the technical scheme also solves the problems of the virtual scene simulation fog technology in the aspects of generating the authenticity and the universality of the image data. Through the combined use of the depth estimation and the optical fogging model, the foggy image can be synthesized in the real scene, and the limitation of the virtual scene is avoided. In addition, the synthetic fog image model also presets a target task to determine a fog visual task matching algorithm according to the characteristics of the original public data set, so that a target fog image data set with the original detailed label is generated, the quality and the application universality of the data set are further improved, and the fog image data set with high quality and the detailed label is successfully constructed. The data set not only can improve the application effect of the image processing and the computer vision model in the foggy environment, but also can provide more real and efficient data support for tasks such as target detection and fusion in the foggy scene.
Drawings
FIG. 1 is a schematic diagram of an atomized image composition framework of a foggy image dataset composition method based on depth estimation in one embodiment;
FIG. 2 is a flow chart of a method for foggy-day image dataset synthesis based on depth estimation in one embodiment;
FIG. 3 is a schematic diagram of a depth estimation model framework based on ResNet in one embodiment;
FIG. 4 is a flow diagram of a standard optical model composite atomized image in an embodiment;
FIG. 5 is a graph showing a comparison of the results of the composition of the thick fog data of MFNet public data sets in a daytime scene according to one example;
FIG. 6 is a graph showing comparison of the results of the composition of the thick fog data of MFNet public data sets in a night scene according to one example;
FIG. 7 is a graph showing a comparison of the results of the composition of the thick fog data of RoadScene public data sets in a daytime scene according to one example;
FIG. 8 is a graph showing comparison of the results of the composition of the thick fog data of RoadScene public data sets in a night scene according to one example;
FIG. 9 is a schematic diagram of data set construction classification in one embodiment;
FIG. 10 is a graphical representation of a comparison of synthetic fog images generated by different fog concentration parameters in one embodiment;
FIG. 11 is a block diagram of a foggy day image data aggregation device based on depth estimation in one embodiment;
fig. 12 is an internal structural diagram of a computer device in one embodiment.
Detailed Description
The present invention will be described in further detail with reference to the drawings and examples, in order to make the objects, technical solutions and advantages of the present invention more apparent. It should be understood that the specific embodiments described herein are for purposes of illustration only and are not intended to limit the scope of the invention.
The foggy day image data set synthesis method based on depth estimation provided by the invention can be applied to a foggy image synthesis frame shown in fig. 1, and comprises three modules, namely a depth estimation model, an optical foggy model and a synthetic foggy data set, and the depth estimation of an image scene, the foggy processing of the optical model and the pairing and arrangement of the synthetic foggy data are respectively realized.
In one embodiment, as shown in fig. 2, there is provided a foggy day image data set synthesizing method based on depth estimation, which is described by taking an example that the method is applied to the foggy image synthesizing frame in fig. 1, and includes the following steps:
and 202, constructing a foggy day image data synthesis model.
The foggy day image data synthesis model comprises a depth estimation model, an optical foggy adding model and a synthetic foggy image model.
And 204, after the depth estimation model is optimally trained by adopting a monocular depth estimation method according to the public data set, inputting the target image scene data into the optimized depth estimation model for semantic deepening, and outputting the synthetic fog effect of the depth map corresponding to the target image scene data.
And 206, generating an atomized image after performing atomization treatment on the depth map through the optical atomization model according to the synthetic fog effect.
And step 208, determining a foggy visual task matching algorithm of the synthetic foggy image model according to the foggy image and a preset target task, and generating a target foggy image data set according to the foggy visual task matching algorithm.
In the foggy day image data set synthesis method based on depth estimation, the optical phenomenon in the foggy day environment can be more accurately simulated by introducing the combination mode of the depth estimation model and the optical foggy model. After the depth estimation model is optimized and trained, the depth information of the target image is obtained by using the model, and semantic deepening is carried out, so that a depth map which is more in line with the real world fog effect is generated. Then, the target image is atomized according to the depth information of the scene through an optical fogging model, and an image with a natural fogging effect is generated. And secondly, the atomization process of the optical atomization model is guided by the depth map information, so that the atomization effect which is finer and adapts to the complex outdoor environment is realized. Specifically, the depth map can provide spatial position information of different objects and backgrounds in the scene, so that the optical fog adding model generates fog layers with different densities and layers according to the actual structure of the scene, and a more real fog effect is simulated. In addition, the technical scheme also solves the problems of the virtual scene simulation fog technology in the aspects of generating the authenticity and the universality of the image data. Through the combined use of the depth estimation and the optical fogging model, the foggy image can be synthesized in the real scene, and the limitation of the virtual scene is avoided. In addition, the synthetic fog image model also presets a target task to determine a fog visual task matching algorithm according to the characteristics of the original public data set, so that a target fog image data set with the original detailed label is generated, the quality and the application universality of the data set are further improved, and the fog image data set with high quality and the detailed label is successfully constructed. The data set not only can improve the application effect of the image processing and the computer vision model in the foggy environment, but also can provide more real and efficient data support for tasks such as target detection and fusion in the foggy scene.
In one embodiment, the depth estimation model is built based on ResNet, wherein the depth estimation model includes a feature fusion block, a residual convolution block, and an adaptive output block.
It is worth noting that a depth estimation model is used to generate a depth map of the target image scene. The depth estimation model is shown in fig. 3. The model is built based on ResNet frames, is formed by combining a feature fusion plate, a residual convolution plate and a self-adaptive output plate, and can realize conversion from an input image to a corresponding depth map. The model is trained using large-scale data of multiple classical datasets using a monocular depth estimation method. In particular, training of each dataset is considered an independent task aimed at seeking an approximate pareto optimal solution on the dataset.
In one embodiment, a monocular depth estimation method is used to define a loss function for each of the training data sets trained on the depth estimation model based on different training data sets in the public data set:
;
;
wherein, the For the size of the training data set i,For scale and translational invariance loss,For a gradient matching item with multiple scales and unchanged scales, M is the number of pixel points contained in the corresponding image of the training data set,As the disparity prediction value, a disparity value,As the true value of the tag is,As a type of loss function,Is a super parameter. And minimizing the loss function by adopting a least square method to obtain a scale and translation invariant loss function:
;
;
where s is a resolution parameter, t is a position parameter, To normalize the disparity prediction value to the label value difference,For the disparity map differences at scale k,To be aligned in the x-axis directionIs used for the gradient of (a),To be aligned in the y-axis directionIs used for the gradient of (a),Is a multi-scale varying scale level.
And enhancing the training task by introducing a 3D data set into each training data set, and taking the minimum value of all the trained scales and translation invariant loss functions as a model optimization parameter if the scale and translation invariant loss functions of the optimized training according to the enhanced task are reduced:
;
wherein, the For the enhanced task of the current training,Is a model parameter. Otherwise, the scale and translation invariable loss function is reserved, and the current trained enhanced task is discarded, so that a task data set to be optimized is obtained. And optimizing the depth estimation model according to the scale and translation invariant loss function and the task data set to be optimized, and training a target image through the optimized depth estimation model.
In one embodiment, inputting target image scene data and the target image into an optimized depth estimation model for semantic deepening, outputting a predicted depth map, and performing brightness inversion on the predicted depth map after gray level conversion to generate a first gray level map;
Extracting scene depth information of the first gray scale map by adopting a standard optical algorithm:
;
;
wherein, the In order to atomize the image,In order to be an image of the object,In the case of a transmission diagram,Is the light of the atmosphere, and is the light of the atmosphere,For a pixel point in the image,For the depth of the scene,Is the atmospheric scattering coefficient. And generating a synthetic fog effect of the depth map corresponding to the target image scene data according to the scene depth information.
It should be noted that, based on the standard optical model, the atomized image is synthesized by the depth map and the original clear image, and the specific process is shown in fig. 4. Firstly, gray level conversion and brightness inversion are carried out on the depth map obtained in the previous stage, and a corresponding gray level map is generated. Subsequently, the optical formula is used:
;
wherein, the In order to atomize the image,In order to be an image of the object,In the case of a transmission diagram,Is the light of the atmosphere, and is the light of the atmosphere,Is a pixel point in the image. The transmission map can be represented by the formulaWhereinThe depth of the scene is represented as,Is the atmospheric scattering coefficient.
In one embodiment, after the visible light image of the depth map is subjected to superposition atomization processing through the optical atomization model according to the synthetic fog effect, an atomization image corresponding to the depth map is generated.
In one embodiment, the target tasks include constructing a foggy visual task matching algorithm performance test dataset, constructing a foggy infrared-visible pixel level pairing fusion dataset, and constructing a foggy target detection dataset.
In one embodiment, a performance test data set of a foggy visual task matching algorithm is constructed, and the method specifically comprises the steps of selecting MFNet semantic segmentation parts of a public data set, and firstly extracting and reconstructing 4-channel data of the semantic segmentation parts to obtain a 3-channel RGB image. Then, the image is classified according to the original classification label into daytime and night scenes. The classified data are input into a depth estimation model, corresponding fog brightness is matched, and fog image data sets with different concentrations in daytime and night are generated by adjusting fog concentration parameters. The data set reserves the original image information and is matched with the depth map, so that the data set is suitable for effect test and contrast evaluation of a foggy day visual task matching algorithm, and subjective visual comparison and objective index evaluation are supported. Examples of partial datasets are shown in fig. 5 and 6.
It is worth to say that, by adopting the method, the synthetic foggy day images with multiple scenes and different concentrations and the corresponding depth map can be quickly generated only by inputting clear images.
In one embodiment, a foggy infrared-visible pixel level paired fusion dataset is constructed by conducting an experiment based on RoadScene public datasets. The dataset independently divides the infrared and visible images. Firstly classifying 221 visible light images, dividing the visible light images into a daytime scene under a strong illumination condition and a night scene under a weak illumination condition, and generating a corresponding depth map by using a depth estimation model. In view of the fact that the daytime scene in the data set is sufficient in sunlight and the night scene is more in light, a higher atomization brightness value is selected to be matched with scene conditions. Finally, the atomized visible light image is paired with the original infrared image, and a data set comprising the visible light image, the infrared image, the depth map and the atomized image which are paired at the pixel level is constructed, so that the method is suitable for an image fusion task. At the same time, preserving the original image facilitates evaluation of fusion algorithm based on subjective vision and objective indexes, such as contrast analysis of indexes of Visual Information Fidelity (VIF) and mutual information quantity (MI). The foggy-day infrared-visible light pixel level pairing fusion dataset constructed based on RoadScene dataset shows that the dataset realizes pairing of pixel levels, namely a visible light image, a depth map, an atomized image and an infrared image from left to right as shown in fig. 7 and 8.
It is worth to say that the synthetic fog superposition is performed on the visible light image based on the existing visible light-infrared paired data set. The investigation shows that the foggy weather has no influence on infrared imaging basically, so that the original infrared image data and labels are reserved, and a visible light-infrared pixel level pairing data set under foggy conditions is constructed. The data set can effectively fill the blank of the foggy-day multi-mode study, and provides a large amount of reliable test data for the related image fusion algorithm.
In one embodiment, a foggy-day target detection dataset is constructed by selecting MFNet detection parts in the public dataset, and classifying the detection parts according to illumination conditions, wherein the detection parts are classified into strong illumination, insufficient illumination and weak illumination as shown in fig. 9. The visible light image data in the categories are input into a depth estimation model, after the corresponding fog brightness is matched, different fog concentration parameters are selected to generate fog day visible light image data with different fog concentrations under the illumination condition shown in fig. 10, and after the fog day visible light image data is paired with an infrared image in the original data set and a detection label, a fog day target detection data set is obtained. Fig. 10 is a partial image data comparison display of the dataset, from top to bottom, of a day fog, a mist scene, a night fog, and a mist scene.
It is worth to say that the fog adding processing is carried out on the existing marked public data set, and a foggy day image with clear marks is generated so as to support training and generalization of a foggy day target detection algorithm. Meanwhile, as the interference of the infrared image on fog is small, the infrared and visible light paired data set is selected and utilized to evaluate the performance of the detection algorithm on different image data, so that the overall accuracy of the detection algorithm is improved.
It should be understood that, although the steps in the flowcharts of fig. 2 and 4 are shown in order as indicated by the arrows, these steps are not necessarily performed in order as indicated by the arrows. The steps are not strictly limited to the order of execution unless explicitly recited herein, and the steps may be executed in other orders. Moreover, at least some of the steps in fig. 2, 4 may comprise a plurality of sub-steps or phases, which are not necessarily performed at the same time, but may be performed at different times, which are not necessarily performed sequentially, but may be performed alternately or alternately with at least some of the other steps or sub-steps of other steps.
In one embodiment, as shown in FIG. 11, there is provided a foggy day image data integration apparatus based on depth estimation, comprising a synthetic model construction module 1102, a synthetic fog effect generation module 1104, a foggy image generation module 1106, and a synthetic fog data module 1108, wherein:
The synthetic model construction module 1102 is configured to construct a foggy day image data synthetic model. The foggy day image data synthesis model comprises a depth estimation model, an optical foggy adding model and a synthetic foggy image model.
The synthetic fog effect generating module 1104 is configured to perform optimization training on the depth estimation model by using a monocular depth estimation method according to the public data set, input the target image scene data to the optimized depth estimation model for semantic deepening, and output the synthetic fog effect of the depth map corresponding to the target image scene data.
And the atomized image generating module 1106 is configured to generate an atomized image after performing atomization processing on the depth map through the optical atomization model according to the synthetic fog effect.
The synthetic fog data module 1108 is configured to determine a foggy day visual task matching algorithm of the synthetic fog image model according to the foggy image and a preset target task, and generate a target foggy day image dataset according to the foggy day visual task matching algorithm.
For a detailed definition of a foggy day image data set synthesizing device based on depth estimation, reference may be made to the definition of a foggy day image data set synthesizing method based on depth estimation hereinabove, and the detailed description thereof will be omitted. The foregoing module of the foggy day image data integration device based on depth estimation may be implemented in whole or in part by software, hardware, or a combination thereof. The above modules may be embedded in hardware or may be independent of a processor in the computer device, or may be stored in software in a memory in the computer device, so that the processor may call and execute operations corresponding to the above modules.
In one embodiment, a computer device is provided, which may be a terminal, and the internal structure thereof may be as shown in fig. 12. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Wherein the processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface of the computer device is used for communicating with an external terminal through a network connection. The computer program, when executed by a processor, implements a foggy day image data set synthesis method based on depth estimation. The display screen of the computer equipment can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer equipment can be a touch layer covered on the display screen, can also be keys, a track ball or a touch pad arranged on the shell of the computer equipment, and can also be an external keyboard, a touch pad or a mouse and the like.
It will be appreciated by persons skilled in the art that the structures shown in fig. 11-12 are block diagrams of only portions of structures associated with the present inventive arrangements and are not limiting of the computer device to which the present inventive arrangements may be implemented, and that a particular computer device may include more or fewer components than shown, or may be combined with certain components, or have different arrangements of components.
In one embodiment, a computer device is provided comprising a memory storing a computer program and a processor that when executing the computer program performs the steps of:
and constructing a foggy day image data synthesis model. The foggy day image data synthesis model comprises a depth estimation model, an optical foggy adding model and a synthetic foggy image model.
And (3) after the depth estimation model is optimized and trained by adopting a monocular depth estimation method according to the public data set, inputting the target image scene data into the optimized depth estimation model for semantic deepening, and outputting the synthetic fog effect of the depth map corresponding to the target image scene data.
And generating an atomized image after performing atomization treatment on the depth map through the optical atomization model according to the synthetic fog effect.
Determining a foggy visual task matching algorithm for synthesizing a foggy image model according to the foggy image and a preset target task, and generating a target foggy image data set according to the foggy visual task matching algorithm
Those skilled in the art will appreciate that implementing all or part of the above described methods may be accomplished by way of a computer program stored on a non-transitory computer readable storage medium, which when executed, may comprise the steps of the embodiments of the methods described above. Any reference to memory, storage, database, or other medium used in embodiments provided herein may include non-volatile and/or volatile memory. The nonvolatile memory can include Read Only Memory (ROM), programmable ROM (PROM), electrically Programmable ROM (EPROM), electrically Erasable Programmable ROM (EEPROM), or flash memory. Volatile memory can include Random Access Memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms such as Static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double Data Rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (SYNCHLINK) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), among others.
The technical features of the above embodiments may be arbitrarily combined, and all possible combinations of the technical features in the above embodiments are not described for brevity of description, however, as long as there is no contradiction between the combinations of the technical features, they should be considered as the scope of the description.
The above examples illustrate only a few embodiments of the invention, which are described in detail and are not to be construed as limiting the scope of the invention. It should be noted that it will be apparent to those skilled in the art that several variations and modifications can be made without departing from the spirit of the invention, which are all within the scope of the invention. Accordingly, the scope of protection of the present invention is to be determined by the appended claims.

Claims (5)

1.一种基于深度估计的雾天图像数据集合成方法,其特征在于,所述方法包括:1. A method for synthesizing foggy image datasets based on depth estimation, comprising: 构建雾天图像数据合成模型;所述雾天图像数据合成模型包括:深度估计模型、光学加雾模型以及合成雾图像模型;Constructing a foggy image data synthesis model; the foggy image data synthesis model includes: a depth estimation model, an optical fogging model, and a synthetic fog image model; 根据公开数据集采用单目深度估计方法对所述深度估计模型进行优化训练后,将目标图像场景数据输入至已优化的深度估计模型进行语义深化,输出所述目标图像场景数据对应深度图的合成雾效应;After optimizing and training the depth estimation model using a monocular depth estimation method based on a public dataset, the target image scene data is input into the optimized depth estimation model for semantic deepening, and a synthetic fog effect corresponding to the depth map of the target image scene data is output; 将目标图像场景数据即所述目标图像输入至已优化的深度估计模型进行语义深化,输出预测深度图,所述预测深度图通过灰度转换后进行亮度反转,生成第一灰度图;Inputting target image scene data, i.e., the target image, into an optimized depth estimation model for semantic deepening, outputting a predicted depth map, and performing brightness inversion on the predicted depth map through grayscale conversion to generate a first grayscale map; 采用标准光学算法提取所述第一灰度图的场景深度信息:The scene depth information of the first grayscale image is extracted using a standard optical algorithm: 其中,为雾化图像,为目标图像,为透射图,为大气光,为图像中的像素点,为场景深度,为大气散射系数;in, For the fog image, is the target image, is the transmission diagram, is atmospheric light, is the pixel in the image, is the scene depth, is the atmospheric scattering coefficient; 根据所述场景深度信息利用专业光学理论生成所述目标图像场景数据对应深度图的合成雾效应;Generate a synthetic fog effect of a depth map corresponding to the target image scene data using professional optical theory according to the scene depth information; 根据所述合成雾效应通过所述光学加雾模型对所述深度图进行雾化处理后,生成雾化图像;After performing a fogging process on the depth map using the optical fogging model according to the synthetic fog effect, a fogged image is generated; 根据所述雾化图像与预设的目标任务确定所述合成雾图像模型的雾天视觉任务匹配算法,根据所述雾天视觉任务匹配算法生成目标雾天图像数据集;所述目标任务包括:构建雾天视觉任务匹配算法性能测试数据集、构建雾天红外-可见光像素级配对融合数据集以及构建雾天目标检测数据集。The foggy visual task matching algorithm of the synthetic foggy image model is determined based on the foggy image and the preset target task, and a target foggy image dataset is generated based on the foggy visual task matching algorithm. The target tasks include: constructing a foggy visual task matching algorithm performance test dataset, constructing a foggy infrared-visible light pixel-level pairing fusion dataset, and constructing a foggy target detection dataset. 2.根据权利要求1所述的一种基于深度估计的雾天图像数据集合成方法,其特征在于,基于ResNet搭建所述深度估计模型,其中,所述深度估计模型包括:特征融合板块、残差卷积板块以及自适应输出板块。2. The method for synthesizing foggy image datasets based on depth estimation according to claim 1, wherein the depth estimation model is built based on ResNet, wherein the depth estimation model includes: a feature fusion module, a residual convolution module, and an adaptive output module. 3.根据权利要求1所述的一种基于深度估计的雾天图像数据集合成方法,其特征在于,根据公开数据集采用单目深度估计方法对所述深度估计模型进行优化训练,包括:3. The method for synthesizing a foggy image dataset based on depth estimation according to claim 1, wherein the depth estimation model is optimized and trained using a monocular depth estimation method based on a public dataset, comprising: 根据公开数据集中不同的训练数据集采用单目深度估计方法定义每一个所述训练数据集在所述深度估计模型训练的损失函数:According to different training data sets in the public data set, a monocular depth estimation method is used to define the loss function of each training data set in the depth estimation model training: 其中,为训练数据集的大小,为尺度与平移不变性损失,为多尺度与尺度不变的梯度匹配项,M为训练数据集对应图像包含的像素点数,为视差预测值,为标签真实值,为损失函数的类型,为超参数;in, For training data set The size of is the scale and translation invariance loss, is the multi-scale and scale-invariant gradient matching term, M is the number of pixels in the corresponding image of the training dataset, is the disparity prediction value, is the true value of the label, is the type of loss function, is a hyperparameter; 采用最小二乘法对所述损失函数进行最小化,得到尺度与平移不变损失函数:The least squares method is used to minimize the loss function to obtain the scale and translation invariant loss function: 其中,s为分辨率参数,t为位置参数,为归一化后视差预测值与标签值差值,为尺度k处的视差图的差异,为在x轴方向上对的梯度,为在y轴方向上对的梯度,为一个多尺度变化的尺度等级;Among them, s is the resolution parameter, t is the position parameter, is the difference between the normalized disparity prediction value and the label value, is the difference of the disparity map at scale k , In the x-axis direction The gradient, In the y-axis direction The gradient, is a scale level with multiple scale changes; 通过对每一个所述训练数据集中引入3D数据集增强训练任务,若根据已增强任务优化训练的所述尺度与平移不变损失函数减少,则取训练后全部尺度与平移不变损失函数的最小值作为模型优化参数:By introducing a 3D dataset enhancement training task into each of the training datasets, if the scale and translation invariance loss function of the enhanced task optimization training is reduced, the minimum value of all scale and translation invariance loss functions after training is taken as the model optimization parameter: 其中,为当前训练的已增强任务,为模型参数;in, is the augmented task for current training, are model parameters; 反之,则保留尺度与平移不变损失函数,丢弃当前训练的已增强任务,得到待优化任务数据集;Otherwise, the scale and translation invariant loss function is retained, the enhanced task currently being trained is discarded, and the task dataset to be optimized is obtained; 根据所述尺度与平移不变损失函数与所述待优化任务数据集优化所述深度估计模型,通过已优化的所述深度估计模型训练目标图像。The depth estimation model is optimized according to the scale and translation invariant loss function and the task dataset to be optimized, and the target image is trained using the optimized depth estimation model. 4.根据权利要求3所述的一种基于深度估计的雾天图像数据集合成方法,其特征在于,根据所述合成雾效应通过所述光学加雾模型对所述深度图进行雾化处理后,生成原始目标图像对应的雾化图像,包括:4. The method for synthesizing foggy image datasets based on depth estimation according to claim 3, wherein the method comprises: performing fogging processing on the depth map using the optical fogging model according to the synthetic fog effect to generate a fogged image corresponding to the original target image, comprising: 根据所述合成雾效应通过所述光学加雾模型对所述深度图的可见光图像进行叠加雾化处理后,生成所述原始目标图像对应的雾化图像。After the visible light image of the depth map is subjected to superimposed fogging processing through the optical fogging model according to the synthetic fog effect, a fogged image corresponding to the original target image is generated. 5.一种基于深度估计的雾天图像数据集合成装置,其特征在于,用于实现权利要求1至4任一项所述的一种基于深度估计的雾天图像数据集合成方法,所述装置包括:5. A device for synthesizing foggy image data based on depth estimation, characterized in that it is used to implement the method for synthesizing foggy image data based on depth estimation according to any one of claims 1 to 4, the device comprising: 合成模型构建模块,用于构建雾天图像数据合成模型;所述雾天图像数据合成模型包括:深度估计模型、光学加雾模型以及合成雾图像模型;A synthesis model construction module is used to construct a foggy image data synthesis model; the foggy image data synthesis model includes: a depth estimation model, an optical fogging model and a synthetic fog image model; 合成雾效应生成模块,用于根据公开数据集采用单目深度估计方法对所述深度估计模型进行优化训练后,将目标图像场景数据输入至已优化的深度估计模型进行语义深化,输出所述目标图像场景数据对应深度图的合成雾效应;A synthetic fog effect generation module is used to optimize and train the depth estimation model using a monocular depth estimation method based on a public dataset, input the target image scene data into the optimized depth estimation model for semantic deepening, and output a synthetic fog effect corresponding to the depth map of the target image scene data; 雾化图像生成模块,用于根据所述合成雾效应通过所述光学加雾模型对所述深度图进行雾化处理后,生成雾化图像;a fog image generation module, configured to generate a fog image by performing fogging processing on the depth map using the optical fog model according to the synthetic fog effect; 合成雾数据模块,用于根据所述雾化图像与预设的目标任务确定所述合成雾图像模型的雾天视觉任务匹配算法,根据所述雾天视觉任务匹配算法生成目标雾天图像数据集。The synthetic fog data module is used to determine the foggy visual task matching algorithm of the synthetic fog image model according to the fogged image and the preset target task, and generate the target foggy image dataset according to the foggy visual task matching algorithm.
CN202510948609.XA 2025-07-10 2025-07-10 A method and device for synthesizing foggy image data sets based on depth estimation Active CN120451727B (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
CN202510948609.XA CN120451727B (en) 2025-07-10 2025-07-10 A method and device for synthesizing foggy image data sets based on depth estimation

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
CN202510948609.XA CN120451727B (en) 2025-07-10 2025-07-10 A method and device for synthesizing foggy image data sets based on depth estimation

Publications (2)

Publication Number Publication Date
CN120451727A CN120451727A (en) 2025-08-08
CN120451727B true CN120451727B (en) 2025-10-17

Family

ID=96620471

Family Applications (1)

Application Number Title Priority Date Filing Date
CN202510948609.XA Active CN120451727B (en) 2025-07-10 2025-07-10 A method and device for synthesizing foggy image data sets based on depth estimation

Country Status (1)

Country Link
CN (1) CN120451727B (en)

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12579804B2 (en) * 2022-06-22 2026-03-17 POSTECH Research and Business Development Foundation Method and device for learning fog-invariant feature
CN116452470B (en) * 2023-06-20 2023-09-15 深圳市欧冶半导体有限公司 Image defogging method and device based on deep learning staged training
CN117078552A (en) * 2023-08-25 2023-11-17 杭州智元研究院有限公司 A fog synthesis method based on scattering model and depth estimation
CN118570449A (en) * 2024-06-04 2024-08-30 桂林电子科技大学 A UAV target detection method in foggy environment based on improved YOLOv8n
CN120182751A (en) * 2025-03-24 2025-06-20 智洋创新科技股份有限公司 Method and device for synthesizing foggy data set based on depth estimation model in power transmission scenario

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
Towards Robust Monocular Depth Estimation:Mixing Datasets for Zero-shot Cross-dataset Transfer;Rene Ranftl et al.;Computer Vision and Pattern Recognition;20200825;1-14 *
基于深度估计的雾天模拟方法;李靓等;激光与光电子学进展;20230608;1010005 *

Also Published As

Publication number Publication date
CN120451727A (en) 2025-08-08

Similar Documents

Publication Publication Date Title
WO2022165809A1 (en) Method and apparatus for training deep learning model
Wen et al. YOFIR: High precise infrared object detection algorithm based on YOLO and FasterNet
Zhang et al. LISU: Low-light indoor scene understanding with joint learning of reflectance restoration
CN116740261B (en) Image reconstruction method and device and training method and device of image reconstruction model
Li et al. GAN-based controllable image data augmentation in low-visibility conditions for improved roadside traffic perception
CN119810394A (en) A data processing method and system based on image acquisition device
CN118941722A (en) Method, system, terminal and medium for three-dimensional morphology recognition and reconstruction of lunar soil particles
Lv et al. Slfusion: a structure-aware infrared and visible image fusion network for low-light scenes
Li et al. Delving deeper into image dehazing: A survey
Li et al. Multi-scale fusion framework via retinex and transmittance optimization for underwater image enhancement
CN115953330B (en) Texture optimization method, device, equipment and storage medium for virtual scene image
Ruan et al. Low-light image enhancement using dual cross attention
US20210224652A1 (en) Methods and systems for performing tasks on media using attribute specific joint learning
Karwowska et al. Image inpainting and digital camouflage: Methods, applications, and perspectives for remote sensing
Wang et al. A multi-scale attentive recurrent network for image dehazing
Mohamed et al. Integrating EnlightenGAN for enhancing car logo detection under challenging lighting conditions
CN120451727B (en) A method and device for synthesizing foggy image data sets based on depth estimation
Liu et al. GloNeRF: Boosting NeRF capabilities and multi-view consistency in low-light environments
Hou et al. Image inpainting via progressive decoder and gradient guidance
CN118115835A (en) Light guide plate defect small sample data expansion method, system, device and storage medium
Weiher Domain adaptation of HDR training data for semantic road scene segmentation by deep learning
Kankanala et al. A Review of Real-Time Single Image Dehazing Architectures for Embedded Vision under Adverse Weather Conditions
Liu et al. Classification guided thick fog removal network for drone imaging: ClassifyCycle
Tsai et al. Advancing underwater image clarity: a GAN-based approach with residual blocks and linear blending
Zhang Low Light Image Enhancement and Saliency Object Detection

Legal Events

Date Code Title Description
PB01 Publication
PB01 Publication
SE01 Entry into force of request for substantive examination
SE01 Entry into force of request for substantive examination
GR01 Patent grant
GR01 Patent grant