To wget mirror a website is to recursively download its public pages and files with GNU wget, then rewrite the links so the copy opens from disk. This guide is for people who want a command in a terminal. You have a public site you own, and you want HTML, CSS, JavaScript, images, and fonts in a folder.
Only export sites you own or are authorized to copy. A public URL is not permission to republish someone else's website.
TL;DR
- A wget mirror website job is a recursive download plus link rewriting. It does not export a CMS or a builder project file.
- Start at the site root of a site you own. Cap depth. Stay on hostnames you control.
- Use recursive download, page requisites, convert-links, adjust-extension, and no-parent together.
- Add wait, a rate limit, and explicit domains. Do not turn on span-hosts unless you also list those domains.
- Recursive wget reads robots.txt. Honor it, especially on a host you do not operate.
- Open the copy locally. Check the homepage, a deep page, a mobile width, and missing images or stylesheets.
- Switch tools when the site is a JS-heavy builder, when you need a clickable preview, or when the owner will not run a terminal.
Owned files are ordinary HTML, CSS, JavaScript, images, and fonts. ChatGPT, Claude, or Cursor can swap text or add a page once the copy is yours.
This is the GNU wget command guide. For HTTrack, SiteSucker, or Cyotek, use the HTTrack alternative guide, the SiteSucker alternative guide, and the Cyotek WebCopy alternative guide. For the broader tool map, read best website downloaders and website copier tools. For the crawl, asset, and deploy checklist, use how to download an entire website. If you mostly want pages that open with no live host, see offline website downloaders.
What wget mirror website means
A wget mirror is a local tree of files that wget fetched by following links. The same command is how you wget download a website, or mirror a website with wget. Start at a URL, walk the public pages, save the files those pages need, and rewrite references so the folder still works after it leaves the original host.
The output is a static snapshot of published pages. The builder project, the CMS database, and the form and store backends stay on the original host. Public presentation can copy. Editing in the original platform, form delivery, carts, memberships, and private dashboards stay put.
GNU wget does this without a GUI. You type a command, you read the log, and you open the folder. If you wanted a desktop copier, you would already be on HTTrack, SiteSucker, or Cyotek WebCopy.
When the command line is enough
Use wget when:
- you can run a terminal and read flag documentation;
- the site serves HTML with normal
<a href>links; - the crawl boundary is a hostname and a directory you understand;
- you will inspect the folder and the log yourself;
- you want a command you can paste into notes or a script.
Skip wget as the first tool when the site loads routes in JavaScript, when images and CSS live on many third-party hosts, or when the person who has to approve the copy will not open a local folder.
A completed wget run is not proof the site copied. The proof is a homepage, an inner page, and a mobile-width window that still load CSS, images, and inner links.
Install wget on Linux, macOS, and Windows
Check first:
wget --version
On many Linux systems wget is already there. If it is missing, install the distro package, for example sudo apt install wget or sudo dnf install wget.
On macOS, install GNU wget with Homebrew, then confirm wget --version again. Apple's default tools do not include it.
On Windows, WSL runs the same flags as Linux. Package managers such as winget or Chocolatey can also install GNU wget. Once the binary prints a version, use the commands below.
Use GNU wget, not a random third-party binary with a similar name.
The core wget mirror command
The command people copy when they want to wget mirror a website looks like this:
wget --mirror --convert-links --adjust-extension --page-requisites --no-parent https://www.example.com/
Replace the URL with a public site you own. Run it from an empty directory so you can see what landed.
--mirror is shorthand for recursive download with infinite depth and timestamping. That is convenient and also easy to over-scope. For a first copy, prefer an explicit recursive download with a depth cap, a domain list, and polite timing:
wget --recursive \
--level=5 \
--page-requisites \
--convert-links \
--adjust-extension \
--no-parent \
--domains=example.com \
--wait=1 \
--random-wait \
--limit-rate=200k \
https://www.example.com/
Start at the site root when you can. --no-parent refuses to climb above the starting directory. If you start at https://www.example.com/blog/ and the CSS lives in /assets/, the mirror will miss those files.
wget writes a folder named after the host, such as www.example.com/. Keep that folder. Do not scatter a second run into a different working directory unless you mean to.
Recursive download, page requisites, and convert-links
These flags do different jobs. Using only one of them is how people end up with HTML and no styles, or with styles and links that still point at the live site.
--recursive / -r. This is the wget recursive download. wget follows links in the HTML it already has, then fetches those URLs, then follows links again, up to --level. It does not run JavaScript. If a menu, gallery, or route exists only after a script runs, recursive wget will not see it.
--page-requisites / -p. wget also fetches the CSS, scripts, images, and other files the saved page references. A recursive crawl without page requisites can produce a skeleton of HTML. Page requisites without recursion is a one-page wget download website. That is useful for a single URL. It does not make a site mirror.
--convert-links / -k. After the crawl finishes, wget rewrites links in the saved HTML and CSS so they point at the local files. Skip this and the copy still borrows from the live host as soon as you open it.
--adjust-extension / -E. wget adds .html or .css when the server served that type without a matching filename. Local browsers and static hosts then open /about as a file instead of a mystery download.
--no-parent / -np. wget stays at or below the starting path. Use it so a blog mirror does not walk the whole marketing site, and so a site-root mirror does not follow ../ into unrelated hosts on a shared path. Pair it with a root URL unless you know the assets live inside that path.
--mirror includes recursion. It does not include page requisites, convert-links, adjust-extension, or no-parent. That is why the common recipe adds those four flags by hand.
Keep the crawl on your domains
By default wget does not span hosts. That is the safe default. A stylesheet or image on cdn.example.com will then be missing unless you allow that host on purpose.
If the public site you own uses a CDN you also control, name both hosts:
wget --recursive --page-requisites --convert-links --adjust-extension --no-parent \
--span-hosts \
--domains=example.com,cdn.example.com \
https://www.example.com/
--span-hosts without --domains is how a wget mirror wanders. A social widget, a font host, or an ad origin can pull wget onto the public web. List every hostname you want. Leave every other hostname out.
Also set a level. --level=5 is a reasonable first cap for a brochure site. Raise it only after you have seen what the first run collected. Infinite depth from --mirror will follow calendars, search pages, filter URLs, and tag combinations until the disk fills.
Reject noisy paths when you know them. Query-string indexes, calendar views, and print URLs are common traps. wget can exclude directories and reject file patterns. Prefer a smaller copy you can check over a huge folder you cannot check.
Wait, rate limits, and robots.txt
A wget recursive download is still a crawler. Treat your own origin with the same manners you would want from anyone else.
--wait=1pauses about one second between requests.--random-waitvaries that pause so the crawl is less bursty.--limit-rate=200kcaps bandwidth so you do not crowd the same server you still use for the live site.
Recursive wget reads robots.txt. Keep that on. If wget skips paths you need on a site you own, change robots.txt so your crawler is allowed, then run again. Do not disable robots handling to copy a site you do not own.
If your own server blocks the default Wget user agent, allow that agent or set --user-agent to a string your server already permits. Do not impersonate a browser to copy a host you do not own.
Logs matter as much as flags. wget prints status codes as it goes. A wall of 404s, 301 loops, or 403s is a signal to stop, tighten --domains and --level, and start over in a clean directory.
How to verify the mirror
Do not trust the last line of the log by itself. Open the files.
- Go into the host folder wget created.
- Serve it locally so root-relative paths resolve. From that folder,
python3 -m http.server 8000is enough, then openhttp://127.0.0.1:8000/in a browser. - Click the homepage, the main nav, the footer, and one deep page such as a blog post, project, or inner landing page.
- Narrow the window to a mobile width. Open the menu. Look for missing icons, overlapping type, or a blank header.
- Watch the browser's network panel with the live site blocked or with the network off. A request that still goes to the original host means convert-links missed it, or the file was never fetched.
- Search the HTML for the live hostname. Leftover
https://www.example.comlinks and CDN URLs are the usual leftovers. - Compare page count with the live nav. If wget saved three HTML files and the live site has a blog index plus twenty posts, the recursive download did not see those routes.
Missing CSS almost always means the stylesheet lived on another host or above --no-parent. Missing images often means srcset, lazy-load attributes, or script-injected URLs. Missing inner pages means the links were never in the HTML wget parsed.
If a page looks fine only while you are online, the copy is still borrowing from the live site. Fix the flags and run again. Do not deploy that folder.
When wget is the wrong tool
wget is the wrong first tool when the public site is not a graph of HTML links.
Hosted builders such as Framer, Webflow, Wix, Squarespace, Shopify themes, Ghost, Elementor, Carrd, Weebly, Duda, and Tilda often load routes, image candidates, and scripts after JavaScript runs. A wget recursive download will save the first HTML and whatever that HTML already names. The rest of the front end can be empty, broken, or still pointed at the builder.
wget is also the wrong handoff when:
- a founder, designer, or client needs to click a working copy in the browser with no local server;
- you want a preview of the copied pages before you keep a ZIP;
- you do not want to own crawl flags, logs, and a local static server.
For those cases, use a GUI copier or a hosted snapshot. HTTrack, SiteSucker, and Cyotek WebCopy still need scope and a local check, but they remove the command line. A browser-based snapshot is stronger when the site is a builder and you need to see the copy first. Export Your Site is one way to create that copy of published pages.
None of these paths move CMS editing, form backends, carts, memberships, or private dashboards. Plan those separately. The public files you do get stay easy to edit.
wget compared with HTTrack, SiteSucker, Cyotek, and hosted snapshots
| Path | Best fit | Strengths | Watch-outs | | --- | --- | --- | --- | | wget | Repeatable CLI mirrors of conventional HTML | Scriptable flags, transparent logs, easy to document | You own every flag, the robots and domain rules, and the local check | | HTTrack | Desktop mirroring with a project UI | Familiar copier, local control, useful logs | Still needs crawl limits and modern-site testing. See the HTTrack alternative guide | | SiteSucker | Mac desktop copying | macOS app, local folder | Mac-focused, same scope and output checks. See the SiteSucker alternative guide | | Cyotek WebCopy | Windows desktop copying | Windows GUI, crawl reports | Windows-only, still a crawler you have to inspect. See the Cyotek WebCopy alternative guide | | Hosted snapshot | Builder sites and a clickable copy | Browser workflow, inspect pages before you keep the files | Static public pages only. Dynamic services stay on the original host |
If wget already produced a folder you can click through offline, another tool may not improve the copy. If wget missed menus, routes, or assets, change tools rather than piling on more flags. The best website downloaders and website copier tools pages map those choices without repeating this command guide.
FAQ
What does wget mirror website mean?
It means using GNU wget to recursively download a public site you own, including the files each page needs, then converting links so the saved folder opens locally. The result is a static copy of published pages. The builder project and the CMS stay on the original host.
How do I wget download a website?
Point GNU wget at the site root with recursive download, page requisites, convert-links, adjust-extension, and no-parent. Add --domains for hostnames you control, a --level cap, and --wait so the crawl stays polite. Use that command only on a site you own or are authorized to copy.
What is a wget recursive download?
A wget recursive download follows links from the starting URL and fetches those pages, up to the depth you set with --level. Combined with --page-requisites, it also saves CSS, scripts, and images those pages name. It does not execute JavaScript, so it cannot discover routes that exist only in a client-side app.
How is --mirror different from --recursive?
--mirror turns on recursive download with infinite depth and timestamping. --recursive is the crawl itself, and you should set --level yourself. Neither flag converts links or fetches page requisites. Add --convert-links, --page-requisites, --adjust-extension, and --no-parent for a usable local site.
How do I keep wget on my domain?
Use --domains with the hostnames you own, keep --no-parent on, and do not pass --span-hosts unless a CDN you control must be included. Then list that CDN in --domains too. Without that pair, span-hosts can follow the public web.
Does wget follow robots.txt?
Yes, in recursive mode wget reads robots.txt and skips disallowed paths. Leave that behavior on. If your own site blocks wget, allow the crawler in robots.txt and run again. Do not turn robots handling off to copy someone else's site.
Can wget copy a JavaScript-heavy builder site?
Usually not well. wget parses HTML and listed assets. It does not run the builder's JavaScript, so generated routes, lazy images, and client-side menus are common misses. For Framer, Webflow, Wix, and similar hosts, test a GUI copier or a hosted snapshot that you can click through.
Is wget better than HTTrack, SiteSucker, or Cyotek WebCopy?
wget is better when you want a documented command and you will verify the folder yourself. HTTrack, SiteSucker, and Cyotek WebCopy are better when you want a desktop UI. None of them remove the need to open the copy, check a deep page, and confirm assets no longer depend on the live host.
How do I know the wget mirror worked?
Serve the host folder locally, open the homepage, click a deep page, resize to mobile, and disconnect the network or block the original host. If styles, images, or inner routes fail, the mirror is incomplete. If links still point at the live site, convert-links did not finish the job.
Can I use wget on any public site?
No. Use wget only on sites you own or are authorized to copy. Public access is not permission to republish, migrate, or keep someone else's pages.
There is a hosted snapshot option for copying published pages.