URL parser
What Is a URL?
A URL (Uniform Resource Locator) is the address system of the web — the closest thing the internet has to a universal postal service — the string that tells a browser exactly where to go and what to ask for. Every link you have ever clicked is a URL, and every URL is a structured instruction with up to six parts, each carrying a distinct piece of meaning. Most people read URLs as opaque strings; developers, marketers, and security analysts read them as parsed data, character by character, because each component answers a different question about where a request goes, what it asks for, and who can see each part along the way.
The anatomy, defined by RFC 3986, breaks down like this:
| Component | Example | What it means |
|---|---|---|
| Scheme | https | The protocol — how to talk to the server |
| Authority | user@example.com:8080 | Who to talk to (and as whom, on which port) |
| Host | example.com | The server's name or IP address |
| Port | 8080 | The door to knock on (default: 443 for HTTPS) |
| Path | /blog/article | Which resource on the server |
| Query | ?page=2&sort=date | Parameters passed to the resource |
| Fragment | #comments | Where to scroll after loading |
The authority deserves unpacking because it is where phishing lives — and where most people's mental model of URLs is wrong. It can contain userinfo (user:password@ — deprecated but still legal, a relic of FTP-era URLs), the host (domain or IP), and an optional port. The host is the only part that determines where your request actually goes — everything before it is decoration, everything after it is instructions for the server. Attackers exploit this with URLs like https://trusted-bank.com@evil.com/login: the real host is evil.com, and trusted-bank.com is just discarded userinfo. The parser exposes this instantly, which is why it pairs naturally with the safe URL checker when a link looks suspicious.
The query string is where the web's business logic hides in plain sight. Everything after the ? is a set of key=value pairs separated by &: UTM tracking parameters (utm_source, utm_medium, utm_campaign), search terms, pagination, session IDs, faceted-navigation filters, and API parameters. Marketers read query strings to audit tracking; developers read them to debug why a page shows the wrong data; privacy-conscious users strip them to remove trackers before sharing links — a ?fbclid=... or ?gclid=... parameter is an ad platform's tracking ID attached to your click, and removing it before forwarding a link is basic link hygiene. The parser splits the query into a readable table so you never have to eyeball-decode ?utm_source=newsletter&utm_medium=email&utm_campaign=spring_sale&discount=20 again.
Then there is percent-encoding, the reason URLs sometimes look like alphabet soup. Characters outside the allowed set (spaces, non-ASCII text, reserved symbols used literally) are encoded as % followed by two hex digits: a space becomes %20, é becomes %C3%A9 (its two UTF-8 bytes, each encoded). Encoding is contextual, which trips people up: a / inside a query value must be encoded as %2F or it becomes a path separator; a ? inside a value must be %3F or it starts a new query string. The parser decodes these back to readable text — and shows you which characters were encoded, which is often the clue in double-encoding bugs — while the URL encoder/decoder handles the reverse trip when you need to build a valid URL from messy input.
Absolute vs. Relative URLs
An absolute URL (https://example.com/blog/post) works everywhere because it contains every component. A relative URL (/blog/post, ../images/logo.png, ?page=2) is a shortcut resolved against the current page's URL — /blog/post means "same scheme, host, and port as where I am now." Relative URLs make sites portable across domains (dev, staging, production) but break the moment content is copied elsewhere — the classic cause of broken images in copied-and-pasted HTML emails. The parser resolves any relative URL against a base you provide, showing exactly where it points.
Punycode and Homograph Attacks
Domain names were originally ASCII-only, but the modern web supports internationalized names — münchen.de, 日本.jp — through an encoding called punycode, which converts Unicode into ASCII prefixed with xn-- (so münchen.de becomes xn--mnchen-3ya.de on the wire). Legitimate and essential for a global internet — and a gift to phishers. A homograph attack registers a domain that looks identical to a trusted one using characters from other alphabets: the Cyrillic "а" instead of the Latin "a" in apple.com renders pixel-identically in most fonts, and the victim cannot tell by looking.
Browsers have partial defenses (displaying punycode when a domain mixes scripts — which is why your browser sometimes shows xn-- gibberish for a link that looked fine in the email), but the reliable defense is inspection, not eyesight: the parser shows every domain in both its visual form and its punycode form, so xn--pple-43d.com cannot hide behind a familiar-looking face. When a link's visual domain and its punycode form tell different stories, you are looking at an attack — no click required. This is the single most valuable habit the parser teaches: never trust what a domain looks like; read what it is.
The Fragment's Secret Life in Single-Page Apps
The fragment was designed for scroll position — #comments jumps the browser to the comments section. But single-page applications (SPAs) repurposed it as a client-side routing mechanism: in a hash-based router, https://app.example.com/#/dashboard/settings never sends /dashboard/settings to the server at all — the server returns the same shell page for every URL, and JavaScript reads the fragment to decide what to render. This made SPAs deployable on any static host, no server rewrite rules needed.
The tradeoffs are real, worth understanding, and a frequent source of subtle bugs. Fragments never reach the server, so server-side analytics, logging, and access control are blind to them — an SPA's "pages" are invisible to the infrastructure. Search engines historically struggled to index fragment-routed content — modern crawlers execute JavaScript, but indexing of fragment-routed pages remains fragile and inconsistent, which is why the industry largely moved to the History API for client-side routing. And the OAuth implicit flow's infamous choice to return access tokens in the fragment (#access_token=...) was because fragments stay client-side — the token never hits server logs — which also made token theft via browser history and referrer leaks a recurring incident class. The parser's fragment display, combined with the knowledge that servers never see it, turns these from mysteries into mechanics.
How to Use the URL Parser
Follow these steps to dissect any URL:
- Paste the URL. Drop in the full address — from an email, a redirect chain, an API log, anywhere. The parser handles absolute URLs directly.
- Add a base URL for relative links. Parsing
/images/logo.png? Enter the page it appeared on so the tool can resolve it to its absolute form. - Read the component breakdown. Scheme, userinfo, host, port, path, query, and fragment appear as labeled fields — no more squinting at a 200-character string.
- Inspect the query parameters. Each
key=valuepair gets its own row, decoded and readable. Spot the UTM tags, the session IDs, the tracking pixels hiding in the URL. - Check the true host. For suspicious links, confirm the host field matches where you expect to go — userinfo tricks and lookalike domains show up here.
- Review the decoded path. Percent-encoded segments render as readable text, revealing what
%E2%9C%93and friends actually say. - Rebuild or normalize. Use the normalized output — lowercase scheme/host, default ports removed, encoding standardized — as the canonical form for comparison or storage.
- Hand off to the right tool. Suspicious host? Run the safe URL checker. Need the server's story? The headers lookup and hosting checker take it from there.
Key Features
| Feature | What it does | Why it matters |
|---|---|---|
| Component breakdown | Splits scheme, host, port, path, query, fragment | Readable structure from opaque strings |
| Query parameter table | Lists every key/value pair decoded | Audits tracking and debugs parameters |
| Percent-decoding | Decodes %XX sequences to readable text | Reveals hidden content in URLs |
| Userinfo exposure | Separates credentials from host | Catches phishing URL tricks |
| Relative URL resolution | Resolves against a base URL | Shows where shortcuts really point |
| URL normalization | Canonical form: case, ports, encoding | Reliable comparison and dedup |
| Punycode display | Shows IDN domains in both forms | Exposes lookalike-domain attacks |
| Private by design | Processed instantly, never stored | Your URLs stay private |
The normalization feature is quieter than it looks but disproportionately useful. HTTPS://Example.COM:443/Path/ and https://example.com/Path/ are the same URL by the spec's rules (scheme and host are case-insensitive; 443 is HTTPS's default port), but they are different strings — so naive string comparison treats them as different pages. Crawlers, analytics, and dedup logic that compare raw strings double-count pages and split link equity, and security allow-lists that compare raw strings can be bypassed by trivial case or encoding variations. Normalizing first — lowercase scheme and host, drop default ports, standardize encoding — gives you the canonical form where equality means equality. Anyone building a crawler, a link database, an analytics pipeline, or an allow-list should normalize every URL at ingestion; the parser shows you exactly what that canonical form looks like.
Use Cases
Marketers Auditing Campaign Tracking
The problem: Your email campaign's links carry UTM parameters — utm_source, utm_medium, utm_campaign — but analytics shows traffic attributed to "direct" instead of the campaign. Somewhere between the email template and the landing page, the tracking broke: a typo'd parameter name, a redirect stripping the query, a copy-paste that truncated the URL.
How this tool helps: Paste the actual link from the sent email and read the query table: every parameter, decoded, in its own row. A missing utm_medium, a utm_source with a typo (newletter instead of newsletter — invisible in reports, devastating to attribution), or a redirect chain (visible when the parsed final URL differs from what you pasted) shows up immediately. It turns "tracking is broken somewhere" into "the redirect from the shortener drops the query string" — a fixable, specific finding you can hand to whoever owns the redirect with the evidence attached.
Developers Debugging APIs and Redirects
The problem: The API returns the wrong data, or the OAuth flow dies with "redirect_uri mismatch," or the webhook payload contains a URL your code cannot handle. The URL looks right at a glance — the bug is in a component you are not examining: double-encoded parameters, a fragment where a query should be, a port that should not be there.
How this tool helps: The parser lays every component bare: you see the double-encoding (%2520 — an encoded percent sign, the classic double-encode tell, meaning something encoded an already-encoded URL), the fragment that never reaches the server (fragments stay client-side — a common OAuth footgun when the provider documents a query parameter and the implementation uses a fragment), the trailing slash that changes routing on strict servers, the case-sensitive path on a Linux host versus the case-insensitive one you tested on Windows. For redirect chains, parsing each hop's Location header URL reveals where parameters get lost. It is the fastest way to stop guessing about a URL and start reading it.
Security Analysts Examining Phishing Links
The problem: A suspicious link arrives: https://secure-login.example-bank.com.verify.tk/signin or https://example-bank.com@192.0.2.44/login. Users see "example-bank" and trust it; analysts need the true host, the punycode behind any internationalized domain, and the full decoded path — in seconds, without clicking.
How this tool helps: The parser extracts the true host (everything the userinfo trick tries to hide), renders punycode lookalikes (xn-- domains) alongside their visual form so homograph attacks are visible, and decodes the path to reveal the actual payload — including the credential-harvesting form fields the path sometimes names. It is the safe, no-click way to answer "where does this really go?" — and the natural first step before the safe URL checker renders its verdict on the destination itself.
SEO Specialists Cleaning Up URL Structures
The problem: The site has accumulated URL cruft: session IDs in query strings, tracking parameters creating infinite duplicate pages, mixed-case paths, trailing-slash inconsistencies. Search engines treat /Page, /page, and /page/ as different URLs, splitting ranking signals across duplicates.
How this tool helps: Parse representative URLs to inventory the mess: which parameters are tracking cruft (candidates for canonical tags or Search Console parameter handling), where case inconsistencies live, which URLs carry session state that should never have been in the address bar — session IDs in URLs leak through referrers, get bookmarked, and get shared, which is why the industry moved them to cookies decades ago. The normalized output shows what the canonical form should be, giving the dev team a concrete spec instead of a vague "clean up the URLs" ticket — and giving you the before/after evidence to prove the cleanup worked.
Frequently Asked Questions
What are the parts of a URL?
A URL has up to six parts: the scheme (https), the authority (userinfo, host, and port), the path (/blog/post), the query string (?key=value), and the fragment (#section). The parser splits any URL into these labeled components.
What is the difference between a URL and a URI?
Every URL is a URI, but not every URI is a URL: URI (Uniform Resource Identifier) is the general concept of an identifier, while URL is the subset that also locates the resource (how and where to retrieve it). A URN like urn:isbn:978-0-13-468599-1 identifies a book without saying where to get it — a URI, but not a URL.
Why do URLs contain %20 and similar codes?
That is percent-encoding: characters outside the URL-legal set are encoded as % plus two hex digits (%20 is a space, %C3%A9 is é). The parser decodes these back to readable text automatically.
What is the query string in a URL?
Everything after the ? — a set of key=value parameters separated by &. It carries tracking tags (UTM), search terms, pagination, and API arguments. The parser lists each parameter in its own decoded row.
Does the fragment (#section) get sent to the server?
No — the fragment is client-side only; browsers strip it before making the request and use it to scroll after loading. Servers never see it, which is why putting data the server needs in the fragment is a classic bug.
How can I tell the real destination of a suspicious link?
Check the host component: it is the only part that determines where the request goes. Watch for userinfo tricks (real-bank.com@evil.com goes to evil.com) and lookalike domains. The parser exposes the true host and renders punycode forms so homograph attacks are visible.
What is URL normalization?
Converting a URL to its canonical form: lowercase scheme and host, remove default ports (443 for HTTPS), standardize percent-encoding. Normalized URLs can be compared reliably — essential for crawlers, analytics, and dedup.
Are the URLs I parse stored?
No. Parsing is processed instantly and never stored — the URL exists only for the duration of your analysis, with no history retained afterward. Paste sensitive links freely; nothing about them outlives the page view.