[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fpy0eVi4CZ4vrKaHxIpn9yF9v2Xu3TWY5r-OAYDW3A2U":3},{"id":4,"question":5,"answer":6,"answerHtml":7,"slug":8,"keywords":9,"article":10,"status":34,"aiModel":39,"aiConfidence":39,"updatedAt":51,"createdAt":51,"_status":50},124,"What is the purpose of using expired domains for C2 servers in penetration testing?","In penetration testing, expired domains that were previously categorized as legitimate by services like Symantec BlueCoat are often chosen as [C2 domains](\u002Fnews\u002Fpenetration-basics-choosing-a-suitable-c2-domain) because they are less likely to be flagged. Tools like CatMyFish automate searching for such domains on expireddomains.net and checking their reputation via sitereview.bluecoat.com.","\u003Cp>In penetration testing, expired domains that were previously categorized as legitimate by services like Symantec BlueCoat are often chosen as [C2 domains](\u002Fnews\u002Fpenetration-basics-choosing-a-suitable-c2-domain) because they are less likely to be flagged. Tools like CatMyFish automate searching for such domains on expireddomains.net and checking their reputation via sitereview.bluecoat.com.\u003C\u002Fp>\u003Cp>\u003Ca href=\"\u002Fnews\u002Fpenetration-basics-choosing-a-suitable-c2-domain\">Read the related One Day Sec article\u003C\u002Fa>\u003C\u002Fp>","what-is-the-purpose-of-using-expired-domains-for-c2-servers-in-penetration-testi-1777485021907","C2 domain, expired domain, penetration testing, Symantec BlueCoat",{"id":11,"title":12,"slug":13,"description":14,"content":15,"contentHtml":30,"cover":31,"author":32,"views":19,"readingTime":33,"status":34,"publishedAt":35,"seo":36,"tags":41,"qaPairs":42,"meta":47,"updatedAt":48,"createdAt":49,"_status":50},33,"Penetration Basics - Choosing a Suitable C2 Domain","penetration-basics-choosing-a-suitable-c2-domain","Learn how to select suitable C2 domains using expireddomains.net, fix CatMyFish bugs, and implement a Python crawler for penetration testing.",{"root":16},{"type":17,"format":18,"indent":19,"version":20,"children":21,"direction":29},"root","",0,1,[22],{"type":23,"format":18,"indent":19,"version":20,"children":24,"direction":29},"paragraph",[25],{"mode":26,"text":27,"type":28,"style":18,"detail":19,"format":19,"version":20},"normal","\u003Chtml>\u003Chead>\u003C\u002Fhead>\u003Cbody>\u003Ch2>0x00 Preface\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>In penetration testing, it is often necessary to select a suitable domain name as a C2 server. So, what kind of domain name can be considered \"suitable\"?\u003C\u002Fp>\u003Cp>expireddomains.net might give you some ideas.\u003C\u002Fp>\u003Cp>Through expireddomains.net, you can query recently expired or deleted domain names, and more importantly, it provides a keyword search function.\u003C\u002Fp>\u003Cp>This article will test the expired domain automation search tool CatMyFish, analyze its principles, fix bugs in it, and use Python to write a crawler to obtain all search results.\u003C\u002Fp>\u003Ch2>0x01 Introduction\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>This article will cover the following:\u003C\u002Fp>\u003Cul>\u003Cli>Testing the expired domain automation search tool CatMyFish\u003C\u002Fli>\u003Cli>Analyzing principles and fixing bugs in CatMyFish\u003C\u002Fli>\u003Cli>Crawler development ideas and implementation details\u003C\u002Fli>\u003Cli>Open-source Python implementation of the crawler code\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>0x02 Testing the Expired Domain Automation Search Tool CatMyFish\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>Download URL:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fgithub.com\u002FMr-Un1k0d3r\u002FCatMyFish\u003C\u002Fp>\u003Ch3>Main Implementation Process\u003C\u002Fh3>\u003Cul>\u003Cli>User inputs keywords\u003C\u002Fli>\u003Cli>The script sends search requests to expireddomains.net for queries\u003C\u002Fli>\u003Cli>Obtains domain list\u003C\u002Fli>\u003Cli>The script sends domains to Symantec BlueCoat for queries\u003C\u002Fli>\u003Cli>Retrieves category for each domain\u003C\u002Fli>\u003C\u002Ful>\u003Cp>expireddomains.net URL:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fwww.expireddomains.net\u002F\u003C\u002Fp>\u003Cp>Symantec BlueCoat URL:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fsitereview.bluecoat.com\u002F\u003C\u002Fp>\u003Ch3>Actual Testing\u003C\u002Fh3>\u003Cp>Requires installation of python library beautifulsoup4\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>pip install beautifulsoup4\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Attempted to search for keyword microsoft, script reported an error as shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019807100_0_fde496a969.jpeg\">\u003C\u002Fp>\u003Cp>The script encountered issues parsing the results\u003C\u002Fp>\u003Cp>Therefore, following the implementation approach of CatMyFish, I wrote my own script for testing\u003C\u002Fp>\u003Cp>Visited expireddomains.net to query the keyword microsoft, code as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>import urllib\u003Cbr>import urllib2\u003Cbr>from bs4 import BeautifulSoup\u003Cbr>url = \"https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?q=microsoft\"\u003Cbr>\u003Cbr>req = urllib2.Request(url)\u003Cbr>res_data = urllib2.urlopen(req)\u003Cbr>\u003Cbr>html = BeautifulSoup(res_data.read(), \"html.parser\")\u003Cbr>\u003Cbr>tds = html.findAll(\"td\", {\"class\": \"field_domain\"})\u003Cbr>\u003Cbr>for td in tds:\u003Cbr>    for a in td.findAll(\"a\", {\"class\": \"namelinks\"}):\u003Cbr>        print a.text\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>A total of 15 results were obtained, as shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019821125_1_3215580c20.jpeg\">\u003C\u002Fp>\u003Cp>Accessing via browser yielded a total of 25 results, as shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019831214_2_67a1d0d810.jpeg\">\u003C\u002Fp>\u003Cp>Comparison revealed that the script obtained fewer results than the browser, likely due to issues in the script's filtering process\u003C\u002Fp>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>Beginners are advised to master the basic usage of beautifulsoup4, which is omitted in this article\u003C\u002Fp>\u003Ch2>0x03 Identifying the cause of the bug\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Ch3>1. Analyze domain tags based on the response to evaluate filtering rules\u003C\u002Fh3>\u003Cp>It is necessary to obtain the received response data and examine the tags corresponding to each domain to determine if there were issues during tag filtering\u003C\u002Fp>\u003Cp>Two methods for viewing response data:\u003C\u002Fp>\u003Ch4>(1) Using the Chrome browser to inspect\u003C\u002Fh4>\u003Cp>F12 -&gt; More tools -&gt; Network conditions\u003C\u002Fp>\u003Cp>Reload the webpage, select ?q=microsoft -&gt; Response\u003C\u002Fp>\u003Cp>As shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019846390_3_d5c3dfdabd.jpeg\">\u003C\u002Fp>\u003Ch4>(2) Using a Python script\u003C\u002Fh4>\u003Cp>The code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>import urllib\u003Cbr>import urllib2\u003Cbr>url = \"https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?q=microsoft\"\u003Cbr>req = urllib2.Request(url)\u003Cbr>res_data = urllib2.urlopen(req)\u003Cbr>print res_data.read()\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Analyzing the response data reveals the cause of the error:\u003C\u002Fp>\u003Cp>Using the original test script can extract the domain names from the following data:\u003C\u002Fp>\u003Cp>\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003C\u002Fp>\u003C\u002Ftd>\u003Ctd class=\"field_domain\">\u003Ca class=\"namelinks\" href=\"\u002Fgoto\u002F1\u002F71h90s\u002F59\u002F?tr=search\" id=\"linksdd-domain71h90s\" rel=\"nofollow\" target=\"_blank\" title=\"MicroSoft.msk.ru\">\u003Cstrong>MicroSoft\u003C\u002Fstrong>.msk.ru\u003C\u002Fa>\u003Cul class=\"kmenucontent\" id=\"links-domain71h90s\" style=\"display:none;\">\u003Cli class=\"first\">\u003Ca class=\"favicons favgodaddy\" href=\"\u002Fgoto\u002F16\u002F75wxyx\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at GoDaddy.com\">GoDaddy.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons favdynadot\" href=\"\u002Fgoto\u002F53\u002F740s95\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at Dynadot.com\">Dynadot.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons favuniregistry\" href=\"\u002Fgoto\u002F66\u002F7252us\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at Uniregistry.com\">Uniregistry.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons favnamecheap\" href=\"\u002Fgoto\u002F43\u002F7459ux\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at Namecheap.com\">Namecheap.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons favonecom\" href=\"\u002Fgoto\u002F57\u002F71gmkr\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at One.com\">One.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons fav123reg\" href=\"\u002Fgoto\u002F48\u002F7254ap\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at 123-reg.co.uk\">123-reg.co.uk\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>However, the response data also contains another type of data:\u003C\u002Fp>\u003Cp>\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003C\u002Fp>\u003C\u002Ftd>\u003Ctd class=\"field_domain\">\u003Ca href=\"\u002Fgoto\u002F1\u002F4o47ng\u002F39\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"NewMicroSoft.com\">New\u003Cstrong>MicroSoft\u003C\u002Fstrong>.com\u003C\u002Fa>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The original test script did not extract the domain information stored in this tag\u003C\u002Fp>\u003Ch2>0x04 Bug Fix\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>Filtering approach:\u003C\u002Fp>\u003Cp>Obtain the content of the first title within the  tag\u003C\u002Fp>\u003Cp>Reason:\u003C\u002Fp>\u003Cp>This allows obtaining domain information stored in both sets of data while filtering out invalid information (such as the domain GoDaddy.com in the second title)\u003C\u002Fp>\u003Cp>Implementation code:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>tds = html.findAll(\"td\", {\"class\": \"field_domain\"})\u003Cbr>for td in tds:\u003Cbr>\tprint td.findAll(\"a\")[0][\"title\"]\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Therefore, the test code to obtain complete query results is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>import urllib\u003Cbr>import urllib2\u003Cbr>import sys\u003Cbr>from bs4 import BeautifulSoup\u003Cbr>\u003Cbr>def SearchExpireddomains(key):\u003Cbr>    url = \"https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?q=\" + key \u003Cbr>    req = urllib2.Request(url)\u003Cbr>    res_data = urllib2.urlopen(req)\u003Cbr>    html = BeautifulSoup(res_data.read(), \"html.parser\")\u003Cbr>    tds = html.findAll(\"td\", {\"class\": \"field_domain\"})\u003Cbr>    for td in tds:\u003Cbr>\tprint td.findAll(\"a\")[0][\"title\"]\u003Cbr>\u003Cbr>if __name__ == \"__main__\":\u003Cbr>    SearchExpireddomains(sys.argv[1])\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Successfully retrieved all results from the first page, test as shown in the image below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019862887_4_482ae48564.jpeg\">\u003C\u002Fp>\u003Ch2>0x05 Get All Query Results\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>expireddomains.net saves 25 results per page. To obtain all results, multiple requests need to be sent to traverse the results across all query pages.\u003C\u002Fp>\u003Cp>First, obtain the total number of all results, then divide by 25 to get the number of pages that need to be queried.\u003C\u002Fp>\u003Ch3>1. Count All Results\u003C\u002Fh3>\u003Cp>Check the Response to find the location indicating the number of search results. The content is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\t\t\u003C\u002Fp>\u003Cdiv class=\"pagescode page_top\">\u003Cbr>\t\t\t\u003Cdiv class=\"addoptions left\">\u003Cbr>\t\t\t\t\t\t\t\t\t\u003Cspan class=\"showfilter\">Show Filter\u003C\u002Fspan>\u003Cbr>\t\t\t\t\u003Cbr>\t\t\t\t\t\t\t\t\t\t\t\t\t\t\u003Cbr>\t\t\t\t\t\t\t\t\t\t\u003Cbr>\t\t\t\t\t\u003Cspan>(About \u003Cstrong>20,213 \u003C\u002Fstrong> Domains)\u003C\u002Fspan>\u003Cp>\u003C\u002Fp>\u003C\u002Fdiv>\u003C\u002Fdiv>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Chrome browser displays as shown in the image below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019867456_5_e8bebd9b62.jpeg\">\u003C\u002Fp>\u003Cp>To simplify the code length, use select() to directly pass in CSS selectors for filtering. After filtering the strong tags, the first tag represents the result count. The corresponding query code is:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>print html.select('strong')[0]\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The output result is \u003Cstrong>20,213 \u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>Extract the numbers from it:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>print html.select('strong')[0].text\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The output result is 20,213\u003C\u002Fp>\u003Cp>Remove the middle \",\":\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>print html.select('strong')[0].text.replace(',', '')\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The output result is 20213\u003C\u002Fp>\u003Cp>Divide by 25 to get the number of pages to query. Note that the string type \"20213\" needs to be converted to integer 20213\u003C\u002Fp>\u003Ch3>2. Guess the query pattern\u003C\u002Fh3>\u003Cp>The query URL for the second page:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?start=25&amp;q=microsoft\u003C\u002Fp>\u003Cp>The query URL for the third page:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?start=50&amp;q=microsoft\u003C\u002Fp>\u003Cp>Find the query pattern. The query URL for the i-th page:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?start=&lt;25*(i-1)&gt;&amp;q=microsoft\u003C\u002Fp>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>Testing shows that expireddomains.net provides a maximum of 550 results for non-logged-in users, across 21 pages.\u003C\u002Fp>\u003Ch3>3. Evaluate the results\u003C\u002Fh3>\u003Cp>In script implementation, it is necessary to evaluate the results: if the results exceed 550, only output 21 pages; if less than 550, output \u003Cresults 25=\"\"> pages.\u003C\u002Fresults>\u003C\u002Fp>\u003Ch3>4. Simulate browser access (alternative)\u003C\u002Fh3>\u003Cp>When using a script to automatically query multiple pages, if the website employs anti-crawling mechanisms, real data cannot be obtained.\u003C\u002Fp>\u003Cp>Testing shows that expireddomains.net has not enabled anti-crawling mechanisms.\u003C\u002Fp>\u003Cp>If in the future, expireddomains.net enables anti-crawling mechanisms, the script needs to simulate browser requests by adding headers such as User-Agent.\u003C\u002Fp>\u003Cp>View the Chrome browser to obtain request information, as shown in the figure below.\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019871173_6_395cf80231.jpeg\">\u003C\u002Fp>\u003Cp>By comparing the request, adding header information can bypass it.\u003C\u002Fp>\u003Cp>Example code:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>req.add_header(\"User-Agent\", \"Mozilla\u002F5.0 (Windows NT 6.1) AppleWebKit\u002F537.36 (KHTML, like Gecko) Chrome\u002F65.0.3325.181 Safari\u002F537.36\")\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Complete code implementation address:\u003C\u002Fp>\u003Cp>An open-source project\u003C\u002Fp>\u003Cp>Actual testing:\u003C\u002Fp>\u003Cp>Search keyword microsoftoffices, results less than 550, as shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019875477_7_00223965a9.jpeg\">\u003C\u002Fp>\u003Cp>Search keyword microsoft, results greater than 550, only 21 pages displayed, as shown in the figure\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fuploads\u002Fdocx_image_1770019879594_8_42360ab8e5.jpeg\">\u003C\u002Fp>\u003Cp>Compared with the content accessed via Web, the results are the same, test successful\u003C\u002Fp>\u003Ch2>0x06 Summary\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>This article tested the expired domain automation search tool CatMyFish, analyzed its principles, fixed bugs in it, used Python to write a crawler to obtain all collection results, shared development ideas, and open-sourced the code.\u003C\u002Fp>\u003C\u002Fbody>\u003C\u002Fhtml>","text","ltr","\u003Chtml>\u003Chead>\u003C\u002Fhead>\u003Cbody>\u003Ch2>0x00 Preface\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>In penetration testing, it is often necessary to select a suitable domain name as a C2 server. So, what kind of domain name can be considered \"suitable\"?\u003C\u002Fp>\u003Cp>expireddomains.net might give you some ideas.\u003C\u002Fp>\u003Cp>Through expireddomains.net, you can query recently expired or deleted domain names, and more importantly, it provides a keyword search function.\u003C\u002Fp>\u003Cp>This article will test the expired domain automation search tool CatMyFish, analyze its principles, fix bugs in it, and use Python to write a crawler to obtain all search results.\u003C\u002Fp>\u003Ch2>0x01 Introduction\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>This article will cover the following:\u003C\u002Fp>\u003Cul>\u003Cli>Testing the expired domain automation search tool CatMyFish\u003C\u002Fli>\u003Cli>Analyzing principles and fixing bugs in CatMyFish\u003C\u002Fli>\u003Cli>Crawler development ideas and implementation details\u003C\u002Fli>\u003Cli>Open-source Python implementation of the crawler code\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>0x02 Testing the Expired Domain Automation Search Tool CatMyFish\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>Download URL:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fgithub.com\u002FMr-Un1k0d3r\u002FCatMyFish\u003C\u002Fp>\u003Ch3>Main Implementation Process\u003C\u002Fh3>\u003Cul>\u003Cli>User inputs keywords\u003C\u002Fli>\u003Cli>The script sends search requests to expireddomains.net for queries\u003C\u002Fli>\u003Cli>Obtains domain list\u003C\u002Fli>\u003Cli>The script sends domains to Symantec BlueCoat for queries\u003C\u002Fli>\u003Cli>Retrieves category for each domain\u003C\u002Fli>\u003C\u002Ful>\u003Cp>expireddomains.net URL:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fwww.expireddomains.net\u002F\u003C\u002Fp>\u003Cp>Symantec BlueCoat URL:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fsitereview.bluecoat.com\u002F\u003C\u002Fp>\u003Ch3>Actual Testing\u003C\u002Fh3>\u003Cp>Requires installation of python library beautifulsoup4\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>pip install beautifulsoup4\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Attempted to search for keyword microsoft, script reported an error as shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019807100_0_fde496a969-1.jpeg\">\u003C\u002Fp>\u003Cp>The script encountered issues parsing the results\u003C\u002Fp>\u003Cp>Therefore, following the implementation approach of CatMyFish, I wrote my own script for testing\u003C\u002Fp>\u003Cp>Visited expireddomains.net to query the keyword microsoft, code as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>import urllib\u003Cbr>import urllib2\u003Cbr>from bs4 import BeautifulSoup\u003Cbr>url = \"https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?q=microsoft\"\u003Cbr>\u003Cbr>req = urllib2.Request(url)\u003Cbr>res_data = urllib2.urlopen(req)\u003Cbr>\u003Cbr>html = BeautifulSoup(res_data.read(), \"html.parser\")\u003Cbr>\u003Cbr>tds = html.findAll(\"td\", {\"class\": \"field_domain\"})\u003Cbr>\u003Cbr>for td in tds:\u003Cbr>    for a in td.findAll(\"a\", {\"class\": \"namelinks\"}):\u003Cbr>        print a.text\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>A total of 15 results were obtained, as shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019821125_1_3215580c20-1.jpeg\">\u003C\u002Fp>\u003Cp>Accessing via browser yielded a total of 25 results, as shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019831214_2_67a1d0d810-1.jpeg\">\u003C\u002Fp>\u003Cp>Comparison revealed that the script obtained fewer results than the browser, likely due to issues in the script's filtering process\u003C\u002Fp>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>Beginners are advised to master the basic usage of beautifulsoup4, which is omitted in this article\u003C\u002Fp>\u003Ch2>0x03 Identifying the cause of the bug\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Ch3>1. Analyze domain tags based on the response to evaluate filtering rules\u003C\u002Fh3>\u003Cp>It is necessary to obtain the received response data and examine the tags corresponding to each domain to determine if there were issues during tag filtering\u003C\u002Fp>\u003Cp>Two methods for viewing response data:\u003C\u002Fp>\u003Ch4>(1) Using the Chrome browser to inspect\u003C\u002Fh4>\u003Cp>F12 -&gt; More tools -&gt; Network conditions\u003C\u002Fp>\u003Cp>Reload the webpage, select ?q=microsoft -&gt; Response\u003C\u002Fp>\u003Cp>As shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019846390_3_d5c3dfdabd-1.jpeg\">\u003C\u002Fp>\u003Ch4>(2) Using a Python script\u003C\u002Fh4>\u003Cp>The code is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>import urllib\u003Cbr>import urllib2\u003Cbr>url = \"https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?q=microsoft\"\u003Cbr>req = urllib2.Request(url)\u003Cbr>res_data = urllib2.urlopen(req)\u003Cbr>print res_data.read()\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Analyzing the response data reveals the cause of the error:\u003C\u002Fp>\u003Cp>Using the original test script can extract the domain names from the following data:\u003C\u002Fp>\u003Cp>\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003C\u002Fp>\u003C\u002Ftd>\u003Ctd class=\"field_domain\">\u003Ca class=\"namelinks\" href=\"\u002Fgoto\u002F1\u002F71h90s\u002F59\u002F?tr=search\" id=\"linksdd-domain71h90s\" rel=\"nofollow\" target=\"_blank\" title=\"MicroSoft.msk.ru\">\u003Cstrong>MicroSoft\u003C\u002Fstrong>.msk.ru\u003C\u002Fa>\u003Cul class=\"kmenucontent\" id=\"links-domain71h90s\" style=\"display:none;\">\u003Cli class=\"first\">\u003Ca class=\"favicons favgodaddy\" href=\"\u002Fgoto\u002F16\u002F75wxyx\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at GoDaddy.com\">GoDaddy.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons favdynadot\" href=\"\u002Fgoto\u002F53\u002F740s95\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at Dynadot.com\">Dynadot.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons favuniregistry\" href=\"\u002Fgoto\u002F66\u002F7252us\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at Uniregistry.com\">Uniregistry.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons favnamecheap\" href=\"\u002Fgoto\u002F43\u002F7459ux\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at Namecheap.com\">Namecheap.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons favonecom\" href=\"\u002Fgoto\u002F57\u002F71gmkr\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at One.com\">One.com\u003C\u002Fa>\u003C\u002Fli>\u003Cli>\u003Ca class=\"favicons fav123reg\" href=\"\u002Fgoto\u002F48\u002F7254ap\u002F59\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"Register at 123-reg.co.uk\">123-reg.co.uk\u003C\u002Fa>\u003C\u002Fli>\u003C\u002Ful>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>However, the response data also contains another type of data:\u003C\u002Fp>\u003Cp>\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\u003C\u002Fp>\u003C\u002Ftd>\u003Ctd class=\"field_domain\">\u003Ca href=\"\u002Fgoto\u002F1\u002F4o47ng\u002F39\u002F?tr=search\" rel=\"nofollow\" target=\"_blank\" title=\"NewMicroSoft.com\">New\u003Cstrong>MicroSoft\u003C\u002Fstrong>.com\u003C\u002Fa>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The original test script did not extract the domain information stored in this tag\u003C\u002Fp>\u003Ch2>0x04 Bug Fix\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>Filtering approach:\u003C\u002Fp>\u003Cp>Obtain the content of the first title within the  tag\u003C\u002Fp>\u003Cp>Reason:\u003C\u002Fp>\u003Cp>This allows obtaining domain information stored in both sets of data while filtering out invalid information (such as the domain GoDaddy.com in the second title)\u003C\u002Fp>\u003Cp>Implementation code:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>tds = html.findAll(\"td\", {\"class\": \"field_domain\"})\u003Cbr>for td in tds:\u003Cbr>\tprint td.findAll(\"a\")[0][\"title\"]\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Therefore, the test code to obtain complete query results is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>import urllib\u003Cbr>import urllib2\u003Cbr>import sys\u003Cbr>from bs4 import BeautifulSoup\u003Cbr>\u003Cbr>def SearchExpireddomains(key):\u003Cbr>    url = \"https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?q=\" + key \u003Cbr>    req = urllib2.Request(url)\u003Cbr>    res_data = urllib2.urlopen(req)\u003Cbr>    html = BeautifulSoup(res_data.read(), \"html.parser\")\u003Cbr>    tds = html.findAll(\"td\", {\"class\": \"field_domain\"})\u003Cbr>    for td in tds:\u003Cbr>\tprint td.findAll(\"a\")[0][\"title\"]\u003Cbr>\u003Cbr>if __name__ == \"__main__\":\u003Cbr>    SearchExpireddomains(sys.argv[1])\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Successfully retrieved all results from the first page, test as shown in the image below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019862887_4_482ae48564-1.jpeg\">\u003C\u002Fp>\u003Ch2>0x05 Get All Query Results\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>expireddomains.net saves 25 results per page. To obtain all results, multiple requests need to be sent to traverse the results across all query pages.\u003C\u002Fp>\u003Cp>First, obtain the total number of all results, then divide by 25 to get the number of pages that need to be queried.\u003C\u002Fp>\u003Ch3>1. Count All Results\u003C\u002Fh3>\u003Cp>Check the Response to find the location indicating the number of search results. The content is as follows:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>\t\t\u003C\u002Fp>\u003Cdiv class=\"pagescode page_top\">\u003Cbr>\t\t\t\u003Cdiv class=\"addoptions left\">\u003Cbr>\t\t\t\t\t\t\t\t\t\u003Cspan class=\"showfilter\">Show Filter\u003C\u002Fspan>\u003Cbr>\t\t\t\t\u003Cbr>\t\t\t\t\t\t\t\t\t\t\t\t\t\t\u003Cbr>\t\t\t\t\t\t\t\t\t\t\u003Cbr>\t\t\t\t\t\u003Cspan>(About \u003Cstrong>20,213 \u003C\u002Fstrong> Domains)\u003C\u002Fspan>\u003Cp>\u003C\u002Fp>\u003C\u002Fdiv>\u003C\u002Fdiv>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Chrome browser displays as shown in the image below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019867456_5_e8bebd9b62-1.jpeg\">\u003C\u002Fp>\u003Cp>To simplify the code length, use select() to directly pass in CSS selectors for filtering. After filtering the strong tags, the first tag represents the result count. The corresponding query code is:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>print html.select('strong')[0]\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The output result is \u003Cstrong>20,213 \u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>Extract the numbers from it:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>print html.select('strong')[0].text\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The output result is 20,213\u003C\u002Fp>\u003Cp>Remove the middle \",\":\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>print html.select('strong')[0].text.replace(',', '')\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>The output result is 20213\u003C\u002Fp>\u003Cp>Divide by 25 to get the number of pages to query. Note that the string type \"20213\" needs to be converted to integer 20213\u003C\u002Fp>\u003Ch3>2. Guess the query pattern\u003C\u002Fh3>\u003Cp>The query URL for the second page:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?start=25&amp;q=microsoft\u003C\u002Fp>\u003Cp>The query URL for the third page:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?start=50&amp;q=microsoft\u003C\u002Fp>\u003Cp>Find the query pattern. The query URL for the i-th page:\u003C\u002Fp>\u003Cp>https:\u002F\u002Fwww.expireddomains.net\u002Fdomain-name-search\u002F?start=&lt;25*(i-1)&gt;&amp;q=microsoft\u003C\u002Fp>\u003Cp>\u003Cstrong>Note:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cp>Testing shows that expireddomains.net provides a maximum of 550 results for non-logged-in users, across 21 pages.\u003C\u002Fp>\u003Ch3>3. Evaluate the results\u003C\u002Fh3>\u003Cp>In script implementation, it is necessary to evaluate the results: if the results exceed 550, only output 21 pages; if less than 550, output \u003Cresults 25=\"\"> pages.\u003C\u002Fresults>\u003C\u002Fp>\u003Ch3>4. Simulate browser access (alternative)\u003C\u002Fh3>\u003Cp>When using a script to automatically query multiple pages, if the website employs anti-crawling mechanisms, real data cannot be obtained.\u003C\u002Fp>\u003Cp>Testing shows that expireddomains.net has not enabled anti-crawling mechanisms.\u003C\u002Fp>\u003Cp>If in the future, expireddomains.net enables anti-crawling mechanisms, the script needs to simulate browser requests by adding headers such as User-Agent.\u003C\u002Fp>\u003Cp>View the Chrome browser to obtain request information, as shown in the figure below.\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019871173_6_395cf80231-1.jpeg\">\u003C\u002Fp>\u003Cp>By comparing the request, adding header information can bypass it.\u003C\u002Fp>\u003Cp>Example code:\u003C\u002Fp>\u003Ctable>\u003Ctbody>\u003Ctr>\u003Ctd>\u003Cp>req.add_header(\"User-Agent\", \"Mozilla\u002F5.0 (Windows NT 6.1) AppleWebKit\u002F537.36 (KHTML, like Gecko) Chrome\u002F65.0.3325.181 Safari\u002F537.36\")\u003C\u002Fp>\u003C\u002Ftd>\u003C\u002Ftr>\u003C\u002Ftbody>\u003C\u002Ftable>\u003Cp>Complete code implementation address:\u003C\u002Fp>\u003Cp>An open-source project\u003C\u002Fp>\u003Cp>Actual testing:\u003C\u002Fp>\u003Cp>Search keyword microsoftoffices, results less than 550, as shown in the figure below\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019875477_7_00223965a9-1.jpeg\">\u003C\u002Fp>\u003Cp>Search keyword microsoft, results greater than 550, only 21 pages displayed, as shown in the figure\u003C\u002Fp>\u003Cp>\u003Cimg alt=\"Alt text\" src=\"\u002Fapi\u002Fmedia\u002Ffile\u002Fdocx_image_1770019879594_8_42360ab8e5-1.jpeg\">\u003C\u002Fp>\u003Cp>Compared with the content accessed via Web, the results are the same, test successful\u003C\u002Fp>\u003Ch2>0x06 Summary\u003C\u002Fh2>\u003Cp>---\u003C\u002Fp>\u003Cp>This article tested the expired domain automation search tool CatMyFish, analyzed its principles, fixed bugs in it, used Python to write a crawler to obtain all collection results, shared development ideas, and open-sourced the code.\u003C\u002Fp>\u003C\u002Fbody>\u003C\u002Fhtml>",1666,"Onedaysec",6,"published","2026-02-02T08:19:47.664Z",{"title":37,"description":14,"keywords":38,"ogImage":39,"canonicalUrl":39,"noIndex":40},"Choosing C2 Domain for Penetration Testing with Expired Domains","penetration testing, C2 domain, expired domains, CatMyFish, Python crawler, bug fix, cybersecurity",null,false,[],{"docs":43,"hasNextPage":40},[44,45,46,4],127,126,125,{"title":39,"description":39,"image":39},"2026-07-24T15:37:15.289Z","2026-07-23T16:01:02.260Z","draft","2026-07-23T16:03:50.122Z"]