[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fzo54dUfJjOm1rc7UAAGIq8ELmMyktOL86czqbPrkJUo":3},{"id":4,"question":5,"answer":6,"answerHtml":7,"slug":8,"keywords":9,"article":10,"status":34,"aiModel":39,"aiConfidence":39,"updatedAt":51,"createdAt":51,"_status":50},765,"How does the parser extract the email subject, sender, and recipients from the XML file?","The parser uses Python's standard library `xml.dom.minidom` to navigate the SOAP XML structure. For fields like subject and body which appear as direct text between tags, it accesses the `firstChild.data` of the corresponding node (e.g., `t:Subject`). For the sender, which is nested under `t:Sender` > `t:Mailbox` > `t:Name`, it retrieves the element by tag name and then gets the child node's data. Recipients and CC are handled via loop extraction over multiple `t:Mailbox` child nodes under `t:ToRecipients` or `t:CcRecipients`. These techniques build on the SOAP message structure introduced in the [Exchange Web Service (EWS) Development Guide 2 – SOAP XML Message](\u002Fnews\u002Fexchange-web-service-ews-development-guide-2-soap-xml-message).","\u003Cp>The parser uses Python&#39;s standard library `xml.dom.minidom` to navigate the SOAP XML structure. For fields like subject and body which appear as direct text between tags, it accesses the `firstChild.data` of the corresponding node (e.g., `t:Subject`). For the sender, which is nested under `t:Sender` &gt; `t:Mailbox` &gt; `t:Name`, it retrieves the element by tag name and then gets the child node&#39;s data. Recipients and CC are handled via loop extraction over multiple `t:Mailbox` child nodes under `t:ToRecipients` or `t:CcRecipients`. These techniques build on the SOAP message structure introduced in the [Exchange Web Service (EWS) Development Guide 2 – SOAP XML Message](\u002Fnews\u002Fexchange-web-service-ews-development-guide-2-soap-xml-message).\u003C\u002Fp>\u003Cp>\u003Ca href=\"\u002Fnews\u002Fexchange-web-service-ews-development-guide-3-soap-xml-parser\">Read the related One Day Sec article\u003C\u002Fa>\u003C\u002Fp>","how-does-the-parser-extract-the-email-subject-sender-and-recipients-from-the-xml-1777481798596","xml.dom.minidom, email extraction, SOAP XML parsing",{"id":11,"title":12,"slug":13,"description":14,"content":15,"contentHtml":30,"cover":31,"author":32,"views":19,"readingTime":33,"status":34,"publishedAt":35,"seo":36,"tags":41,"qaPairs":42,"meta":47,"updatedAt":48,"createdAt":49,"_status":50},188,"Exchange Web Service (EWS) Development Guide 3 – SOAP XML Parser","exchange-web-service-ews-development-guide-3-soap-xml-parser","Learn to build a SOAP XML parser for EWS to automatically extract email details like subject, sender, body, and attachments using Python's standard libraries.",{"root":16},{"type":17,"format":18,"indent":19,"version":20,"children":21,"direction":29},"root","",0,1,[22],{"type":23,"format":18,"indent":19,"version":20,"children":24,"direction":29},"paragraph",[25],{"mode":26,"text":27,"type":28,"style":18,"detail":19,"format":19,"version":20},"normal","\u003Chtml>\u003Chead>\u003C\u002Fhead>\u003Cbody>\u003Ch2>0x00 Preface\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>In the previous article \"Exchange Web Service (EWS) Development Guide 2 – SOAP XML message\", the use of SOAP XML messages was introduced, demonstrating how to access Exchange resources using hash via Python.\u003C\u002Fp>\u003Cp>When reading emails through SOAP XML messages, we often encounter the following issue: since each email corresponds to a raw XML file containing complete email information, manually analyzing emails consumes significant effort.\u003C\u002Fp>\u003Cp>Therefore, this article will introduce an implementation method for a SOAP XML parser, developing a tool to automatically extract valuable email information and improve reading efficiency.\u003C\u002Fp>\u003Ch2>0x01 Introduction\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>This article will cover the following:\u003C\u002Fp>\u003Cul>\u003Cli>Applicable Environment\u003C\u002Fli>\u003Cli>Design Approach\u003C\u002Fli>\u003Cli>Open-source Python Implementation Code\u003C\u002Fli>\u003Cli>Code Development Details\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>0x02 Design Approach\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>To read all emails in the inbox via SOAP XML messages, the following steps are required:\u003C\u002Fp>\u003Col>\u003Cli>Use the listmailofinbox command of ewsManage.py to obtain the ItemId and ChangeKey for each email\u003C\u002Fli>\u003Cli>Iteratively use the getmail command of ewsManage.py, passing in the ItemId and ChangeKey corresponding to each email\u003C\u002Fli>\u003Cli>Save the returned results as XML format files separately, with each XML file corresponding to one email\u003C\u002Fli>\u003C\u002Fol>\u003Cp>To ensure the versatility of the SOAP XML parser and its compatibility with different tools, the SOAP XML parser is designed with a file manager structure. Selecting an XML file will automatically trigger parsing, extract valuable information, and display it. The design follows these principles:\u003C\u002Fp>\u003Cul>\u003Cli>The development language is Python, and to enhance convenience, only Python's standard libraries are used\u003C\u002Fli>\u003Cli>The file manager involves Python GUI development, using the standard GUI library Tkinter\u003C\u002Fli>\u003Cli>The SOAP (Simple Object Access Protocol) is essentially an XML protocol, and parsing uses the standard library xml.dom.minidom\u003C\u002Fli>\u003C\u002Ful>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>If parsing XML files using string matching, escape characters must also be considered\u003C\u002Fp>\u003Ch2>0x03 Program Implementation\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Ch3>1. Implementation of the File Manager\u003C\u002Fh3>\u003Cp>Using Tkinter:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fdocs.python.org\u002F3\u002Flibrary\u002Ftk.html\u003C\u002Fp>\u003Cp>Secondary development can be based on the open-source file-manager-mask, with the following modifications:\u003C\u002Fp>\u003Cul>\u003Cli>Remove the image display functionality\u003C\u002Fli>\u003Cli>Remove the text editing functionality\u003C\u002Fli>\u003Cli>Add XML file parsing functionality\u003C\u002Fli>\u003C\u002Ful>\u003Ch3>2. XML File Parsing\u003C\u002Fh3>\u003Cp>Usage of xml.dom.minidom:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fdocs.python.org\u002F3\u002Flibrary\u002Fxml.dom.minidom.html\u003C\u002Fp>\u003Cp>The following content needs to be extracted here:\u003C\u002Fp>\u003Cul>\u003Cli>Email subject\u003C\u002Fli>\u003Cli>Sender\u003C\u002Fli>\u003Cli>Recipient\u003C\u002Fli>\u003Cli>CC (Carbon Copy)\u003C\u002Fli>\u003Cli>Receipt time\u003C\u002Fli>\u003Cli>Attachment name\u003C\u002Fli>\u003Cli>Body content\u003C\u002Fli>\u003C\u002Ful>\u003Cp>In data extraction, there are the following different scenarios:\u003C\u002Fp>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>XML tags are case-sensitive\u003C\u002Fp>\u003Ch4>(1) Extracting node attributes\u003C\u002Fh4>\u003Cp>Example format of the response message:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Cm:getitemresponsemessage responseclass=\"Success\">\u003C\u002Fm:getitemresponsemessage>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Extract the attribute \"ResponseClass\" of the node \"m:GetItemResponseMessage\", the code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_response = dom.getElementsByTagName(\"m:GetItemResponseMessage\")\u003Cbr>print(data_response[0].getAttribute(\"ResponseClass\"))\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch4>(2) Directly extracting data between tag pairs\u003C\u002Fh4>\u003Cp>Example format of the email subject:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:subject>123\u003C\u002Ft:subject>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Example format of the body content:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:body bodytype=\"Text\" istruncated=\"false\">123\u003C\u002Ft:body>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Example format of received time:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:datetimereceived>2021-01-11T11:08:50Z\u003C\u002Ft:datetimereceived>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>To extract the content of node \"t:Subject\", the code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_subject = dom.getElementsByTagName(\"t:Subject\")\u003Cbr>print(data_subject[0].firstChild.data)\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Example format of sender:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:sender>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test1\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test1@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\u003C\u002Ft:sender>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Consider parent and child nodes here\u003C\u002Fp>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>There is usually only one sender, so no need to consider loop extraction\u003C\u002Fp>\u003Cp>Extract the content of child node \"t:Name\" under parent node \"t:Sender\", code as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_from = dom.getElementsByTagName(\"t:Sender\")\u003Cbr>print(data_from[0].getElementsByTagName(\"t:Name\")[0].firstChild.data)\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch4>(3) Loop extraction of data between tag pairs\u003C\u002Fh4>\u003Cp>Recipient format example:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:torecipients>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test2\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test2@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test3\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test3@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\u003C\u002Ft:torecipients>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Example format for CC recipients:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:ccrecipients>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test2\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test2@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test3\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test3@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\u003C\u002Ft:ccrecipients>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Attachment format example:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:attachments>\u003Cbr>\t\u003Ct:fileattachment>\u003Cbr>\t\t\u003Ct:attachmentid id=\"AAMk**1\">\u003Cbr>\t\t\u003Ct:name>image1.jpg\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:contenttype>image\u002Fjpeg\u003C\u002Ft:contenttype>\u003Cbr>\t\t\u003Ct:contentid>image1.jpg@11111111.11111111\u003C\u002Ft:contentid>\u003Cbr>\t\t\u003Ct:size>1024\u003C\u002Ft:size>\u003Cbr>\t\t\u003Ct:lastmodifiedtime>2021-01-01T01:01:01\u003C\u002Ft:lastmodifiedtime>\u003Cbr>\t\t\u003Ct:isinline>true\u003C\u002Ft:isinline>\u003Cbr>\t\t\u003Ct:iscontactphoto>false\u003C\u002Ft:iscontactphoto>\u003Cbr>\t\u003C\u002Ft:attachmentid>\u003C\u002Ft:fileattachment>\u003Cbr>\t\u003Ct:fileattachment>\u003Cbr>\t\t\u003Ct:attachmentid id=\"AAMk**2\">\u003Cbr>\t\t\u003Ct:name>image2.jpg\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:contenttype>image\u002Fjpeg\u003C\u002Ft:contenttype>\u003Cbr>\t\t\u003Ct:contentid>image2.jpg@11111111.11111112\u003C\u002Ft:contentid>\u003Cbr>\t\t\u003Ct:size>1024\u003C\u002Ft:size>\u003Cbr>\t\t\u003Ct:lastmodifiedtime>2021-01-01T01:01:01\u003C\u002Ft:lastmodifiedtime>\u003Cbr>\t\t\u003Ct:isinline>true\u003C\u002Ft:isinline>\u003Cbr>\t\t\u003Ct:iscontactphoto>false\u003C\u002Ft:iscontactphoto>\u003Cbr>\t\t\u003C\u002Ft:attachmentid>\u003C\u002Ft:fileattachment>\u003Cbr>\t\u003C\u002Ft:attachments>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Here we need to consider parent nodes and sibling nodes\u003C\u002Fp>\u003Cp>Extract the content of all child nodes \"t:Name\" under the parent node \"t:ToRecipients\", the code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_to = dom.getElementsByTagName(\"t:ToRecipients\")\u003Cbr>data_to_name = data_to[0].getElementsByTagName(\"t:Name\")\u003Cbr>for i in range(len(data_to_name)):\u003Cbr>\tprint(data_to_name[i].firstChild.data)\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The above code skips the judgment of the node \"t:Mailbox\". If we add the judgment, the code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_to = dom.getElementsByTagName(\"t:ToRecipients\")\u003Cbr>data_to_mailbox = data_to[0].getElementsByTagName(\"t:Mailbox\")\u003Cbr>for i in range(len(data_to_mailbox)):\u003Cbr>\tprint(data_to_mailbox[i].getElementsByTagName(\"t:Name\")[0].firstChild.data)\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>After completing data extraction from the XML file, consider how to display the data in the file manager window\u003C\u002Fp>\u003Cp>The insert function will be used here\u003C\u002Fp>\u003Cp>Parameter description:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fdocs.python.org\u002F3.8\u002Flibrary\u002Ftkinter.ttk.html?highlight=insert#tkinter.ttk.Notebook.insert\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>insert(pos, child, **kw)\u003Cbr>Inserts a pane at the specified position.\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>For the pos parameter, END represents insertion from the last line, while a number represents insertion from a specified line (e.g., 1.0 for the first line)\u003C\u002Fp>\u003Cp>The complete code has been uploaded to GitHub at the following address:\u003C\u002Fp>\u003Cp>An open-source project\u003C\u002Fp>\u003Cp>The code supports the following features:\u003C\u002Fp>\u003Cul>\u003Cli>File manager for viewing multiple files, allowing file switching via keyboard arrow keys\u003C\u002Fli>\u003Cli>XML file parsing, capable of automatically extracting valuable information from Exchange SOAP XML messages and flagging XML files that do not conform to the format\u003C\u002Fli>\u003C\u002Ful>\u003Cp>The running interface is shown in the figure below:\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770017239659_0_04a279d688.jpeg\">\u003C\u002Fp>\u003Cp>Subsequently, a complete Exchange GUI client program can be developed by integrating ewsManage.py to enable reading Exchange emails using hashes\u003C\u002Fp>\u003Ch2>0x04 Summary\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>This article introduces an implementation method for a SOAP XML parser, detailing the development of a tool to automatically extract email information from Exchange SOAP XML messages, including open-source Python implementation code and an analysis of code development specifics\u003C\u002Fp>\u003C\u002Fbody>\u003C\u002Fhtml>","text","ltr","\u003Chtml>\u003Chead>\u003C\u002Fhead>\u003Cbody>\u003Ch2>0x00 Preface\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>In the previous article \"Exchange Web Service (EWS) Development Guide 2 – SOAP XML message\", the use of SOAP XML messages was introduced, demonstrating how to access Exchange resources using hash via Python.\u003C\u002Fp>\u003Cp>When reading emails through SOAP XML messages, we often encounter the following issue: since each email corresponds to a raw XML file containing complete email information, manually analyzing emails consumes significant effort.\u003C\u002Fp>\u003Cp>Therefore, this article will introduce an implementation method for a SOAP XML parser, developing a tool to automatically extract valuable email information and improve reading efficiency.\u003C\u002Fp>\u003Ch2>0x01 Introduction\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>This article will cover the following:\u003C\u002Fp>\u003Cul>\u003Cli>Applicable Environment\u003C\u002Fli>\u003Cli>Design Approach\u003C\u002Fli>\u003Cli>Open-source Python Implementation Code\u003C\u002Fli>\u003Cli>Code Development Details\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>0x02 Design Approach\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>To read all emails in the inbox via SOAP XML messages, the following steps are required:\u003C\u002Fp>\u003Col>\u003Cli>Use the listmailofinbox command of ewsManage.py to obtain the ItemId and ChangeKey for each email\u003C\u002Fli>\u003Cli>Iteratively use the getmail command of ewsManage.py, passing in the ItemId and ChangeKey corresponding to each email\u003C\u002Fli>\u003Cli>Save the returned results as XML format files separately, with each XML file corresponding to one email\u003C\u002Fli>\u003C\u002Fol>\u003Cp>To ensure the versatility of the SOAP XML parser and its compatibility with different tools, the SOAP XML parser is designed with a file manager structure. Selecting an XML file will automatically trigger parsing, extract valuable information, and display it. The design follows these principles:\u003C\u002Fp>\u003Cul>\u003Cli>The development language is Python, and to enhance convenience, only Python's standard libraries are used\u003C\u002Fli>\u003Cli>The file manager involves Python GUI development, using the standard GUI library Tkinter\u003C\u002Fli>\u003Cli>The SOAP (Simple Object Access Protocol) is essentially an XML protocol, and parsing uses the standard library xml.dom.minidom\u003C\u002Fli>\u003C\u002Ful>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>If parsing XML files using string matching, escape characters must also be considered\u003C\u002Fp>\u003Ch2>0x03 Program Implementation\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Ch3>1. Implementation of the File Manager\u003C\u002Fh3>\u003Cp>Using Tkinter:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fdocs.python.org\u002F3\u002Flibrary\u002Ftk.html\u003C\u002Fp>\u003Cp>Secondary development can be based on the open-source file-manager-mask, with the following modifications:\u003C\u002Fp>\u003Cul>\u003Cli>Remove the image display functionality\u003C\u002Fli>\u003Cli>Remove the text editing functionality\u003C\u002Fli>\u003Cli>Add XML file parsing functionality\u003C\u002Fli>\u003C\u002Ful>\u003Ch3>2. XML File Parsing\u003C\u002Fh3>\u003Cp>Usage of xml.dom.minidom:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fdocs.python.org\u002F3\u002Flibrary\u002Fxml.dom.minidom.html\u003C\u002Fp>\u003Cp>The following content needs to be extracted here:\u003C\u002Fp>\u003Cul>\u003Cli>Email subject\u003C\u002Fli>\u003Cli>Sender\u003C\u002Fli>\u003Cli>Recipient\u003C\u002Fli>\u003Cli>CC (Carbon Copy)\u003C\u002Fli>\u003Cli>Receipt time\u003C\u002Fli>\u003Cli>Attachment name\u003C\u002Fli>\u003Cli>Body content\u003C\u002Fli>\u003C\u002Ful>\u003Cp>In data extraction, there are the following different scenarios:\u003C\u002Fp>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>XML tags are case-sensitive\u003C\u002Fp>\u003Ch4>(1) Extracting node attributes\u003C\u002Fh4>\u003Cp>Example format of the response message:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Cm:getitemresponsemessage responseclass=\"Success\">\u003C\u002Fm:getitemresponsemessage>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Extract the attribute \"ResponseClass\" of the node \"m:GetItemResponseMessage\", the code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_response = dom.getElementsByTagName(\"m:GetItemResponseMessage\")\u003Cbr>print(data_response[0].getAttribute(\"ResponseClass\"))\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch4>(2) Directly extracting data between tag pairs\u003C\u002Fh4>\u003Cp>Example format of the email subject:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:subject>123\u003C\u002Ft:subject>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Example format of the body content:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:body bodytype=\"Text\" istruncated=\"false\">123\u003C\u002Ft:body>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Example format of received time:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:datetimereceived>2021-01-11T11:08:50Z\u003C\u002Ft:datetimereceived>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>To extract the content of node \"t:Subject\", the code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_subject = dom.getElementsByTagName(\"t:Subject\")\u003Cbr>print(data_subject[0].firstChild.data)\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Example format of sender:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:sender>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test1\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test1@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\u003C\u002Ft:sender>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Consider parent and child nodes here\u003C\u002Fp>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>There is usually only one sender, so no need to consider loop extraction\u003C\u002Fp>\u003Cp>Extract the content of child node \"t:Name\" under parent node \"t:Sender\", code as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_from = dom.getElementsByTagName(\"t:Sender\")\u003Cbr>print(data_from[0].getElementsByTagName(\"t:Name\")[0].firstChild.data)\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Ch4>(3) Loop extraction of data between tag pairs\u003C\u002Fh4>\u003Cp>Recipient format example:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:torecipients>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test2\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test2@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test3\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test3@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\u003C\u002Ft:torecipients>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Example format for CC recipients:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:ccrecipients>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test2\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test2@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\t\u003Ct:mailbox>\u003Cbr>\t\t\u003Ct:name>test3\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:emailaddress>test3@test.com\u003C\u002Ft:emailaddress>\u003Cbr>\t\t\u003Ct:routingtype>SMTP\u003C\u002Ft:routingtype>\u003Cbr>\t\t\u003Ct:mailboxtype>Mailbox\u003C\u002Ft:mailboxtype>\u003Cbr>\t\u003C\u002Ft:mailbox>\u003Cbr>\u003C\u002Ft:ccrecipients>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Attachment format example:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003Ct:attachments>\u003Cbr>\t\u003Ct:fileattachment>\u003Cbr>\t\t\u003Ct:attachmentid id=\"AAMk**1\">\u003Cbr>\t\t\u003Ct:name>image1.jpg\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:contenttype>image\u002Fjpeg\u003C\u002Ft:contenttype>\u003Cbr>\t\t\u003Ct:contentid>image1.jpg@11111111.11111111\u003C\u002Ft:contentid>\u003Cbr>\t\t\u003Ct:size>1024\u003C\u002Ft:size>\u003Cbr>\t\t\u003Ct:lastmodifiedtime>2021-01-01T01:01:01\u003C\u002Ft:lastmodifiedtime>\u003Cbr>\t\t\u003Ct:isinline>true\u003C\u002Ft:isinline>\u003Cbr>\t\t\u003Ct:iscontactphoto>false\u003C\u002Ft:iscontactphoto>\u003Cbr>\t\u003C\u002Ft:attachmentid>\u003C\u002Ft:fileattachment>\u003Cbr>\t\u003Ct:fileattachment>\u003Cbr>\t\t\u003Ct:attachmentid id=\"AAMk**2\">\u003Cbr>\t\t\u003Ct:name>image2.jpg\u003C\u002Ft:name>\u003Cbr>\t\t\u003Ct:contenttype>image\u002Fjpeg\u003C\u002Ft:contenttype>\u003Cbr>\t\t\u003Ct:contentid>image2.jpg@11111111.11111112\u003C\u002Ft:contentid>\u003Cbr>\t\t\u003Ct:size>1024\u003C\u002Ft:size>\u003Cbr>\t\t\u003Ct:lastmodifiedtime>2021-01-01T01:01:01\u003C\u002Ft:lastmodifiedtime>\u003Cbr>\t\t\u003Ct:isinline>true\u003C\u002Ft:isinline>\u003Cbr>\t\t\u003Ct:iscontactphoto>false\u003C\u002Ft:iscontactphoto>\u003Cbr>\t\t\u003C\u002Ft:attachmentid>\u003C\u002Ft:fileattachment>\u003Cbr>\t\u003C\u002Ft:attachments>\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Here we need to consider parent nodes and sibling nodes\u003C\u002Fp>\u003Cp>Extract the content of all child nodes \"t:Name\" under the parent node \"t:ToRecipients\", the code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_to = dom.getElementsByTagName(\"t:ToRecipients\")\u003Cbr>data_to_name = data_to[0].getElementsByTagName(\"t:Name\")\u003Cbr>for i in range(len(data_to_name)):\u003Cbr>\tprint(data_to_name[i].firstChild.data)\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The above code skips the judgment of the node \"t:Mailbox\". If we add the judgment, the code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>from xml.dom import minidom\u003Cbr>dom = minidom.parse(\"TestMail.xml\")\u003Cbr>data_to = dom.getElementsByTagName(\"t:ToRecipients\")\u003Cbr>data_to_mailbox = data_to[0].getElementsByTagName(\"t:Mailbox\")\u003Cbr>for i in range(len(data_to_mailbox)):\u003Cbr>\tprint(data_to_mailbox[i].getElementsByTagName(\"t:Name\")[0].firstChild.data)\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>After completing data extraction from the XML file, consider how to display the data in the file manager window\u003C\u002Fp>\u003Cp>The insert function will be used here\u003C\u002Fp>\u003Cp>Parameter description:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fdocs.python.org\u002F3.8\u002Flibrary\u002Ftkinter.ttk.html?highlight=insert#tkinter.ttk.Notebook.insert\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>insert(pos, child, **kw)\u003Cbr>Inserts a pane at the specified position.\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>For the pos parameter, END represents insertion from the last line, while a number represents insertion from a specified line (e.g., 1.0 for the first line)\u003C\u002Fp>\u003Cp>The complete code has been uploaded to GitHub at the following address:\u003C\u002Fp>\u003Cp>An open-source project\u003C\u002Fp>\u003Cp>The code supports the following features:\u003C\u002Fp>\u003Cul>\u003Cli>File manager for viewing multiple files, allowing file switching via keyboard arrow keys\u003C\u002Fli>\u003Cli>XML file parsing, capable of automatically extracting valuable information from Exchange SOAP XML messages and flagging XML files that do not conform to the format\u003C\u002Fli>\u003C\u002Ful>\u003Cp>The running interface is shown in the figure below:\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770017239659_0_04a279d688-1.jpeg\">\u003C\u002Fp>\u003Cp>Subsequently, a complete Exchange GUI client program can be developed by integrating ewsManage.py to enable reading Exchange emails using hashes\u003C\u002Fp>\u003Ch2>0x04 Summary\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>This article introduces an implementation method for a SOAP XML parser, detailing the development of a tool to automatically extract email information from Exchange SOAP XML messages, including open-source Python implementation code and an analysis of code development specifics\u003C\u002Fp>\u003C\u002Fbody>\u003C\u002Fhtml>",762,"Onedaysec",4,"published","2026-02-02T07:38:21.199Z",{"title":37,"description":14,"keywords":38,"ogImage":39,"canonicalUrl":39,"noIndex":40},"EWS SOAP XML Parser Guide: Extract Email Data with Python","Exchange Web Services, EWS, SOAP XML parser, Python email parsing, XML data extraction, Tkinter file manager, email automation",null,false,[],{"docs":43,"hasNextPage":40},[44,45,4,46],767,766,764,{"title":39,"description":39,"image":39},"2026-07-24T15:37:11.566Z","2026-07-23T16:02:04.076Z","draft","2026-07-23T16:14:36.881Z"]