Font Size: a A A

Research On Text Steganalysis Techniques Based Markup Language

Posted on:2013-08-05Degree:MasterType:Thesis
Country:ChinaCandidate:Z SunFull Text:PDF
GTID:2248330395476450Subject:Communication and Information System
Abstract/Summary:
With the Continuous development and wide application for information technique, information security has been absorbing more and more attention by people. Text steganography is an information security technique, the secret information is embedded in text carrier, and transmitted though text carrier’s transfer. AS promising technique against text steganography, the primary purpose of text steganalysis is to discover the hidden information, and then extract, recover, destruct the hidden information. The study of steganography technique can prevent criminals disrupt national security and social stability. On the other hand, it also can promote the development of steganography technique.After analyzing the structure of news RSS document, a new information hiding method based on news RSS document is developed. Because the order of the news items does not affect the reading of RSS document, the permutation and combination of<item> tags modules can be used to hide information. Through the rational combination of other two selected information hiding methods based on XML documents, an information hiding system using multi-methods based on news RSS document is constructed. The experimental results show that three kinds of information hiding methods do not conflict. The system has good performances in capacity, invisibility and robustness. This system can be used to embed watermark in RSS document and covert communication.Steganography based on the uppercase-lowercase conversion of the tag letters is a common Markup language information hiding algorithm. After deep research on the steganography, the increased randomness of uppercase-lowercase is found when the secret messages are embedded. The paper describes the randomness using chi-square of tag, and two chi-square detection methods are respectively proposed for sequentially and randomly embedding algorithms. For sequentially embedding algorithm, the tags are extracted in turn, chi-square of tag is computed repeatedly, finally whether the document contain hidden information is determined by the value of P, and the length of hidden information is estimated. The experimental results show that the method can effectively determine whether the documents contain hidden information, and the error of estimation in length is about20bit. For randomly embedding algorithm, all of the tags are extracted to compute the chi-square of tag, using correcting the chi-square of tag. And whether the documents contain hidden information is determined by the value of P. The experimental results show that the method has high detection rate. For natural documents, the false positive is0.0067. And for documents with low embedding rate, it also has a very good detection results. When the embedding rate is greater than10%, the false negative is0.
Keywords/Search Tags:steganalysis, eXtensible Markkup Language(XML), tag, chi-square detection, randomness
Related items