URL extractor
What Is a URL Extractor? Finding Every Link Hidden in Text
A URL extractor scans any block of text and pulls out every web address it contains — http and https links, www addresses without a protocol, links buried in HTML anchor tags, URLs inside CSS and JavaScript, shortened links, and even bare domains mentioned in prose. Paste in a document, a page's HTML source, a chat export, a spreadsheet column, or a log file, and the tool returns a clean list of just the links, one per line, ready to check, archive, or process.
URLs are trickier to detect than they look. A robust extractor has to handle the full anatomy defined by RFC 3986: scheme, authority (with optional userinfo, host, and port), path, query string, and fragment. Real-world text makes it harder — links wrapped in parentheses in Markdown, trailing punctuation that isn't part of the address ("visit example.com."), URLs split across line breaks in emails, HTML entities like & standing in for ampersands, percent-encoded characters, internationalized domain names in Punycode, and the endless variety of URL shorteners. A good extractor resolves these correctly: it strips sentence punctuation from the end of a match, decodes entities, and optionally normalizes the results.
The use cases cluster around auditing and research. SEO specialists paste a page's HTML to inventory every outbound link. Developers scan logs to find every endpoint their app actually called. Researchers collect citations from a pile of papers. Archivists extract links from old documents before the domains die. Moderators pull links from user reports to check them against blocklists. And anyone migrating a website extracts every internal link from the old pages so none get lost in the move. In each case the alternative — eyeballing text and copying links one at a time — is slow, and humans reliably miss links hiding in attributes, scripts, and footnotes.
URL extraction pairs naturally with sibling utilities. If your source is a web page and you also need the contact addresses on it, run the same text through our email extractor afterward. If the extracted list has repeats from navigation menus and footers, our duplicate lines remover collapses them to a unique set in one click. And if the page's markup itself needs slimming down, our HTML minifier strips the whitespace and comments without touching the links.
One caution before you start: extracted URLs can be dangerous. Link lists from untrusted sources — phishing reports, user-generated content, old archives — may contain malicious destinations. Never click through an extracted list blindly; preview or scan suspicious links with a URL reputation checker first. Our tool only lists what it finds — it never fetches the destinations — and your input is processed instantly and never stored.
How to Use the URL Extractor
- Collect your source text. Copy the content containing the links: a document, spreadsheet range, chat log, log file, or a web page's HTML source (right-click → View Page Source, then select all and copy).
- Paste it into the input area. The tool accepts plain text, Markdown, and raw HTML — markup is parsed, not treated as noise.
- Run the extraction. The scanner identifies every URL-shaped string, strips trailing punctuation that belongs to the sentence rather than the link, and decodes HTML entities.
- Choose your output options. Toggle deduplication (recommended when scanning full page source, where navigation links repeat dozens of times), choose whether to include bare domains without a protocol, and decide if you want relative links resolved or listed as-is.
- Filter by type if needed. Narrow the results to a single domain, a file extension (e.g., only .pdf links), or external vs. internal links, depending on your audit.
- Copy or download the list. Copy the results as plain lines or download as .txt/.csv, then feed them into a link checker, spreadsheet, or archiving tool.
- Verify before acting. Extraction finds addresses that look like URLs; it doesn't confirm they resolve. Run important lists through a broken-link checker before publishing or migrating.
Key Features
| Feature | What It Does | Why It Matters |
|---|---|---|
| RFC 3986-aware matching | Parses scheme, host, port, path, query, fragment correctly | Handles complex real-world URLs, not just simple ones |
| HTML source parsing | Reads href, src, action, and other link attributes | Finds links invisible in rendered text — images, scripts, forms |
| Trailing-punctuation trimming | Strips sentence punctuation from link ends | "example.com." in prose extracts as example.com |
| Entity decoding | Converts & and friends back to real characters | Links copied from page source work when pasted |
| Automatic deduplication | Each unique URL listed once | Menus and footers don't flood the results |
| Domain and extension filters | Keep only matching links | Isolate PDFs, one domain, or external links only |
| Shortened-link preservation | Lists short links as found, flags them | You see what's really in the text before expanding |
| Export options | Copy or download as TXT/CSV | Drops into spreadsheets and audit tools directly |
What Counts as a URL? The Anatomy the Extractor Understands
To appreciate what the extractor is doing, it helps to see a URL the way the parser does. Take https://user@example.com:8080/path/page.html?search=test&lang=en#section: the scheme (https) says how to fetch it; the authority (user@example.com:8080) says where, including optional credentials and port; the path (/path/page.html) says what; the query (?search=test&lang=en) carries parameters; and the fragment (#section) points within the document. A naive pattern that only matches "http://something" misses bare www. addresses, protocol-relative links (//cdn.example.com/lib.js), and addresses with unusual but legal characters.
Percent-encoding is the other half of correctness. Spaces and special characters in URLs appear as %20-style sequences, and non-Latin domain names appear as Punycode (xn-- prefixes). The extractor recognizes these as part of the address rather than breaking the match at the first odd character. It also handles the classic prose problem: in "See https://example.com/guide (it's excellent).", the closing parenthesis belongs to the sentence, not the link — except when the URL itself legitimately contains parentheses, as Wikipedia links famously do. Good extractors use balanced-parenthesis heuristics for exactly this case.
What the extractor deliberately does not do is fetch anything. It never resolves short links, never follows redirects, and never checks whether a link is alive — that's the job of a link checker or archiver. Extraction is the inventory step: it tells you what's referenced so downstream tools can verify, archive, or rewrite.
Use Cases
For SEO specialists: "The problem:" auditing every outbound link on a page
The problem: A client page links out to dozens of resources accumulated over years. Some point to dead domains, some to competitors you'd rather not endorse, and some use redirect chains that leak authority. Finding them by reading the page is hopeless — half the links are in the footer, the sidebar, and embedded widgets.
How this tool helps: Paste the page's full HTML source into the extractor and get every link in one list — including the ones in scripts and widgets you'd never see by reading. Filter to external domains, dedupe, and hand the list to a link checker. It's the fastest possible inventory pass, and it catches links that visual audits always miss. Run the page's CSS through our CSS minifier separately if the audit also covers page weight.
For developers: "The problem:" finding every endpoint in a log dump
The problem: A production incident left you with a multi-megabyte log file. You need the distinct set of URLs the application actually requested — to reproduce the failure, to audit third-party calls, or to build an allowlist.
How this tool helps: Paste the log excerpt and get the unique URLs instantly, query strings included. Filter by domain to separate your own API calls from third-party services. It's dramatically faster than crafting the perfect grep, and the dedupe means you see the shape of the traffic instead of a wall of repeated lines.
For researchers: "The problem:" collecting citations from a document pile
The problem: Your literature review spans forty papers and reports, each citing web sources in footnotes, reference lists, and inline prose with inconsistent formatting. You need every cited URL in one list for archiving before link rot sets in.
How this tool helps: Paste each document's text and accumulate the links. The extractor handles the inconsistent formatting — parenthesized links, "available at:" prefixes, broken-across-lines addresses — that make manual collection so error-prone. Dedupe across the whole set, then submit the survivors to the Wayback Machine so your citations survive even when the originals don't.
For site migrators: "The problem:" no link left behind in a redesign
The problem: You're moving a site to a new platform. Every internal link, every PDF, every image reference in the old HTML needs to be accounted for — miss one and you ship broken links on day one.
How this tool helps: Extract links from each old page's source, filter by the old domain, and you have a complete inventory to map against the new URL structure. Pair it with the HTML minifier when optimizing the new templates, and use the extractor again after launch to verify the new pages link only where they should.
For moderators: "The problem:" triaging links in user reports
The problem: Users report suspicious posts as copied text or screenshots-turned-text. You need the actual URLs to check against threat feeds — quickly, without clicking anything.
How this tool helps: Paste the report and get every link as plain text you can inspect safely. The extractor never fetches destinations, so there's no risk of drive-by exposure from the extraction itself. Feed the results to your URL reputation scanner and act on the verdicts.
Link Rot: Why Extracting Links Is an Archival Survival Skill
The average web page has a measurable half-life. Studies of link rot consistently find that a significant percentage of links cited in academic papers, news articles, and documentation stop working within a few years — domains expire, sites restructure, companies get acquired and their blogs vanish. Every link you don't inventory is a link you can't preserve. Extraction is the first step of a simple archival workflow: extract every link from your important documents, verify which ones still resolve, and archive the survivors in the Wayback Machine or a similar service.
This matters most for content with a long shelf life. A technical tutorial that links to library documentation, a thesis with a hundred web citations, a company knowledge base full of vendor links — all of these decay silently. Nobody notices a dead link until a reader clicks it, and by then the original content may be unrecoverable. A yearly routine of extracting and checking links turns link rot from an embarrassing discovery into a maintenance task. For documentation teams, it's worth pairing the link audit with a markup cleanup: our HTML minifier keeps the pages lean while the extractor keeps their references honest.
There's a second, subtler reason to inventory links: redirect chains. A link that technically "works" may bounce through three or four redirects — a shortener, a tracking redirect, a domain migration — before landing. Each hop slows the click, leaks referrer data, and adds a failure point: if any hop in the chain dies, the link dies. Extraction gives you the raw list; a link checker reveals the chains; then you rewrite the worst offenders to point at final destinations. Sites that do this annually feel noticeably snappier to users, because fewer clicks take the scenic route.
Common Extraction Mistakes and How to Avoid Them
Most extraction disappointments come from the input, not the tool. The classic mistake is pasting rendered text instead of source: copying from a rendered web page gives you anchor text ("click here") while the actual URLs stay behind in the HTML. When you need the real destinations, always extract from View Source. Conversely, pasting raw HTML when you only wanted the visible links floods results with stylesheet, script, and tracking-pixel URLs — use the attribute or domain filters to narrow down.
Another frequent gotcha is line-broken URLs. Email clients and PDFs wrap long addresses across lines, sometimes inserting hyphens or spaces at the break. The extractor repairs simple wraps, but a URL broken mid-query-string with an inserted hyphen may extract as two fragments. When a result looks truncated, search the source for the fragment — the remainder is usually on the next line.
Finally, don't confuse extraction with validation, and don't treat the output as a sitemap. An extracted list reflects what the text references, which includes outdated links, example domains (example.com in documentation), and links inside code comments. Before feeding results into any automated process — a crawler, a migration script, a blocklist — skim for these artifacts. Ten minutes of review saves hours of debugging a pipeline that chased documentation examples.
Frequently Asked Questions
Does the extractor visit or check the links it finds?
No. It purely parses text — it never makes network requests, never resolves short links, and never follows redirects. Use a dedicated link checker or uptime tool to verify that extracted URLs actually work.
Can it extract links from a PDF?
Copy the PDF's text and paste it in — links in native-text PDFs extract cleanly. For scanned PDFs, run OCR first. Note that some PDFs store links as annotations rather than visible text; copying the text layer still usually captures them.
Will it find links hidden in HTML attributes like src and action?
Yes. The extractor parses common link-bearing attributes — href, src, srcset, action, cite, data attributes that look like URLs — so images, scripts, stylesheets, and form targets all show up, not just clickable anchor text.
How does it handle shortened links like bit.ly addresses?
It lists them exactly as they appear in the text. It does not expand them, because expansion requires a network request. If you need the destinations, paste the short links into a URL expander service — preferably one that shows the destination without redirecting your browser straight there.
Can I extract only links from a specific domain?
Yes — the domain filter keeps only URLs matching the domain you enter (including subdomains, optionally). This is the fastest way to inventory one site's links inside a mixed document, or to pull all references to a single source from research notes.
What about relative links like /about or ../images/logo.png?
The extractor lists relative references as found, since it can't know the base URL from text alone. If you're scanning a site's HTML, note the site's domain separately; most link checkers accept a base URL to resolve relatives against.
Why do some extracted links have & in them?
That's an HTML entity from page source: in HTML, ampersands in URLs are written as &. The extractor's entity-decoding option converts these back to plain & so the copied link works when pasted into a browser.
Is there a limit to how much text I can paste?
The tool handles typical workloads comfortably — full page sources, long documents, and sizable log excerpts. For truly enormous inputs (tens of megabytes of logs), paste in chunks; browsers themselves get sluggish with gigantic clipboard operations regardless of the tool.
Can it extract email addresses too?
It focuses on URLs, though mailto: links will surface as links. For a proper address list — with dedupe, domain filtering, and de-obfuscation — use our dedicated email extractor instead.
Is my pasted text stored anywhere?
No. Input is processed instantly to produce the link list and is never stored. The tool makes no network requests to the extracted destinations either — your source text never leaves the processing step.