Font Size: a A A

High-fidelity Colorization Of Image Generation

Posted on:2024-08-14Degree:MasterType:Thesis
Country:ChinaCandidate:J M SunFull Text:PDF
GTID:2568306914459894Subject:Information and Communication Engineering
Abstract/Summary:
High-fidelity image colorization is an important part in the generation of Internet content,which aims at convert the input grayscale image into a colorful image,owning wide application prospect in artistic creation and scene restoration.With the rapid development of deep learning,using artificial intelligence technology to generate content and using machine-assisted image creation have become research hotspots,and the image generation effect has also been greatly improved.Image colorization can be divided into two categories:text-guided colorization and fully automatic colorization.(1)Textguided colorization means the colorization guided by text,semantic segmentation masks and other modality inputs.Text-guided colorization is faced with the following research difficulties:Firstly,colorization results are hard to keep consistent with input texts.Due to the gap between the two modalities of text and image,it is difficult to establish an accurate text-image mapping,resulting in the discrepancy between generation results and input texts.Secondly,it is hard to achieve efficient fusion and decoupling of multiple modal inputs.Because semantic segmentation masks from other modalities are also needed to help colorize the specified area,the existing multimodal methods are difficult to achieve high-quality results.(2)Fully automatic colorization is trained in a self-supervised manner without additional user inputs,which simplifies user inputs.However,it appears unreasonable semantic colorization and low saturation due to the multi-modality of automatic colorization,which reduces the image generation quality.For the questions and challenges above,high-fidelity image colorization is a worthy and promising topic.This paper focuses on two directions of text-guided image colorization and fully automatic colorization.The main research of the paper is as follows:Firstly,the paper proposes a vertically structured cross-modal similarity model to establish the association between visual image and textual descriptions for the discrepancy between input texts and colorization results in textguided colorization.Comparing to the existing feature matching technology,we focus on fine-grained modeling,optimize the coarse-grained matching process,and improve the consistency between colorization results and input text.This part serves as a pretraining process for text-guided image colorization.The experimental results show our colorization results are more consistent with the semantics of input text,laying the foundation for the subsequent complete text-guided image colorization framework.Secondly,for efficient multimodal fusion and decoupling in text-guided colorization,the paper proposes novel multi-modality fusion module,which realize conditions interaction and decoupling inspired by batch normalization.Additionally,for high fidelity generation and harmonious composition,this paper proposes a unified framework to remove inter-stage dependencies and avoid the accumulation of inter-stage errors.Experimental results show the method improves image generation quality and makes the overall tone of results more harmonious.Finally,the paper proposes a Transformer-based colorization network to solve the incorrect and unsaturation colorization in automatic colorization.Benefiting from the remote dependency advantage of Transformer,and the grayscale selection module designed,model can reduce the range of color selection to improve the rationality of colorization results.Meanwhile,the paper introduces the color tokens,and use the classification loss constraint for training to increase the saturation and color richness,and further obtain high-fidelity colorized images.The experimental results show that the proposed method can achieve more accurate and reasonable colorization results with higher saturation.
Keywords/Search Tags:image colorization, cross-modal feature matching, multi-modal condition fusion
Related items