R-CLO
CloudScraper
Spiders a target website and regex-scans every page it finds for references to S3, Azure Blob, or DigitalOcean Spaces resources.
OVERVIEW
CloudScraper (github.com/jordanpotti/CloudScraper) is a Python reconnaissance tool that takes a different angle from keyword-permutation brute-forcers: instead of guessing bucket names, it spiders a target website's own pages and regex-scans the raw page source — not just parsed HTML links — for references to cloud storage resources such as S3 buckets (`s3.amazonaws.com`), Azure Blob Storage (`windows.net`), and DigitalOcean Spaces.
Because it matches against raw source rather than only `<a href>` tags, it catches storage URLs that are buried in inline JavaScript, JSON blobs, or comments — places a target's own front-end code references a bucket that a link-only crawler would miss. It supports configurable crawl depth, concurrent worker processes, and scanning a list of targets in one run.
USE CASES
Practical use cases
- 01
Discovering cloud storage buckets referenced inside a target's HTML, inline scripts, or bundled JavaScript.
- 02
Complementing keyword-permutation tools (cloud_enum, CloudBrute) with an evidence-based, site-driven discovery angle.
- 03
Batch-scanning a list of target domains from a client-provided scope file in one pass.
- 04
Tuning crawl depth and process count to balance coverage against scan time on large sites.
QUICK START
During web-focused recon, to find cloud storage links that a target's own site leaks in HTML, inline scripts, or JavaScript bundles.
- Clone the repository and install dependencies (`requests`, `beautifulsoup4`, `termcolor`, `rfc3987`) with pip.
- Confirm the target site is in scope, since this tool actively crawls and downloads every page it can reach.
- Run against a single URL with `-u`, or pass a file of targets with `-l` for a batch scan.
- Adjust `-d` (crawl depth) and `-p` (parallel processes) to fit the size of the target site.
- Review flagged cloud-storage links and verify actual bucket/container accessibility separately.
python3 CloudScraper.py -u https://www.target.com -d 5 -p 4 -vBEFORE YOU RUN IT
What to check before running it
This is an active web spider, not a passive OSINT source — full-site crawling can generate a large amount of traffic and hit rate limits or WAF rules.
Only crawl domains and pages within the authorized engagement scope; a spider will follow links wherever they lead unless bounded.
A regex match for a storage URL is a lead, not a finding — confirm the resource actually exists and is accessible before reporting it.