URL Encoding Explained: What Percent-Encoding Is and How to Encode URLs Correctly

29 September, 2026 • 10 views • 11 minutes read

What URL encoding (percent-encoding) is, which characters must be encoded, the whole-URL mistake to avoid, and how to encode query parameters correctly.

You have seen URLs like https://example.com/search?q=tea%20and%20coffee and probably wondered what that %20 is doing there. That is URL encoding — also called percent-encoding — and it is the mechanism that lets any text travel safely inside a web address. Get it wrong and links break, form data arrives corrupted, and APIs return cryptic errors.

Encode URLs correctly in one click: the free URL encoder handles every character.

This guide explains what URL encoding is, which characters must be encoded, the difference between encoding a full URL and encoding a single value, and how to do it correctly with a free URL encoder.

Why URLs Can't Contain Just Any Character

A URL is a tightly structured string: scheme, host, path, query, fragment. Certain characters are structural — they have jobs. A ? starts the query string, & separates parameters, / separates path segments, # starts the fragment. If you want to use one of those characters as plain data (for example, searching for "fish & chips"), the URL parser cannot tell your data apart from its own syntax.

The second problem is the character set: URLs are defined over ASCII. Non-ASCII characters — accented letters, Cyrillic, Chinese, emoji — simply do not exist in a raw URL. URL encoding solves both problems in one move: it converts unsafe characters into a safe %HH triplet, where HH is the hexadecimal value of the character's byte.

How Percent-Encoding Works

The rule is simple: take each byte of the character (in UTF-8 for anything outside ASCII) and write it as a percent sign followed by two uppercase or lowercase hex digits.

  • A space becomes %20 (byte 0x20).
  • & becomes %26, = becomes %3D, ? becomes %3F, # becomes %23.
  • The letter é becomes %C3%A9 — two bytes in UTF-8, each encoded.
  • An emoji like 😀 becomes four triplets: %F0%9F%98%80.

Decoders reverse the process. A free URL decoder turns %20 back into a space and %C3%A9 back into é — useful when you receive an encoded URL and want to read the original text.

Which Characters Must Be Encoded?

Not every character needs encoding. The safe set — characters that can appear raw in a URL — is small and worth memorising:

  • Unreserved (never encode): A-Z, a-z, 0-9, and - _ . ~
  • Reserved (encode when used as data): ! * ' ( ) ; : @ & = + $ , / ? # [ ]
  • Everything else: spaces, non-ASCII characters, control characters, < > " { } | \ ^ \` — always encode.

A common rule of thumb: if a character is not alphanumeric and not - _ . ~, encode it when it appears inside a value. When in doubt, encode.

The Classic Mistake: Encoding the Whole URL

The most frequent URL-encoding error is encoding the entire URL including its structural characters. If you encode https://example.com/path?q=a&b=c wholesale, the ://, /, ?, and & all become %3A, %2F, %3F, %26 — and the URL stops working entirely.

The correct approach: encode only the values, never the structure. In q=tea & coffee, encode just the value: q=tea%20%26%20coffee. Build the URL first, encode the pieces, then assemble. Every major language's URL library works this way — encodeURIComponent() in JavaScript encodes a value, while encodeURI() leaves the structure intact. Mixing them up is one of the top sources of broken links in web development.

Encoding Query Parameters the Right Way

Query strings are where encoding matters most, because that is where user input lands. Follow these rules:

  1. Encode keys and values separately. Never concatenate raw user input into a URL.
  2. Encode the & and = inside values. A value like a=b must become a%3Db or it will be parsed as two parameters.
  3. Watch out for +. In the application/x-www-form-urlencoded format used by HTML forms, + means a space. So encoding a literal plus sign (in "C++") must produce %2B, not +. This form-encoding quirk causes endless confusion — a plain percent-encoder gives you %20 for spaces, which is correct for URLs; form bodies use + instead.
  4. Encode # in values. A raw # starts the fragment, and everything after it never reaches the server. A value containing # that is not encoded will be silently truncated.

URL Encoding vs Base64 vs HTML Entities

People often confuse three different encodings. They solve different problems:

  • URL encoding (percent-encoding): makes text safe to carry *inside a URL*. Use it when building links, query strings, and redirects.
  • Base64: packs arbitrary binary data into a URL-safe-ish ASCII string for *transmission or storage* — email attachments, embedded images, tokens. Note that standard Base64 output contains +, /, and = which still need URL encoding if placed in a URL; that is why JWTs use the Base64url variant (- and _ instead). A free Base64 encoder handles the conversion step for you.
  • HTML entities: make text safe to render *inside an HTML page* (&, <). They do nothing for URLs.

Using the wrong one is a real bug class: Base64-encoding a URL does not make it clickable-safe, and URL-encoding data that goes into a database achieves nothing. Match the encoding to the context.

Where Encoding Matters in Practice

  • Search and filter links: every user-typed search term in a URL must be encoded. "black & white photos" becomes black%20%26%20white%20photos.
  • Share links and QR codes: QR codes routinely encode full URLs — a QR code generator builds the encoded address for you, but understanding what is inside helps when a scanned code fails.
  • Redirect URLs: a URL passed as a parameter (like ?next=/dashboard) must itself be encoded, or its ? and & will corrupt the outer URL.
  • API calls: REST APIs take parameters in URLs; unencoded values are the single most common cause of "400 Bad Request" errors in integrations.
  • Internationalised domain names: non-ASCII domains are converted to Punycode (münchen.de → xn--mnchen-3ya.de) — a related but separate system from percent-encoding.

How to Verify Your Encoded URLs

A quick verification routine saves hours of debugging:

  1. Decode it back. Paste the encoded URL into a URL decoder and confirm you get the original text. If decoding does not reproduce your input, something was double-encoded or partially encoded.
  2. Watch for double encoding. %2520 means the % itself got encoded — the telltale sign that a value was encoded twice. Decode once more to recover it.
  3. Test in the browser address bar. Paste the full URL and check that the server receives the right parameters. Browser dev tools (Network tab) show exactly what was sent.
  4. Check # handling. If a parameter's value seems truncated, look for an unencoded #.

Encoding Paths, Fragments, and Usernames in URLs

Query parameters get all the attention, but other URL parts need encoding too:

  • Path segments: a filename like my report (final).pdf must be encoded as my%20report%20(final).pdf. Slashes between segments stay raw; slashes *inside* a segment value become %2F.
  • Fragments: anything after # is handled by the browser, not sent to the server, but a fragment containing spaces or # itself still needs encoding to parse correctly.
  • Credentials in URLs: the deprecated user:password@host format breaks the moment a password contains @ or : — which is one more reason modern systems moved credentials out of URLs entirely.

URL Encoding and SEO

Search engines treat ?q=a%20b and ?q=a+b as equivalent in most cases, so encoding choices rarely hurt rankings — but sloppy encoding can:

  • Duplicate content: if the same page is reachable as ?q=café (raw), ?q=caf%C3%A9 (encoded), and ?q=caf%E9 (wrongly encoded Latin-1), search engines may see three URLs. Canonical tags solve this, but consistent encoding prevents it.
  • Broken crawl links: a link with an unencoded space may be "fixed" differently by different crawlers, creating phantom 404s in Search Console.
  • Readability: Google displays decoded URLs in search results where possible, so clean, properly encoded URLs look better to users too.

URL Encoding in Code: The One-Liners That Matter

Almost every language ships correct encoders — use them instead of hand-rolling replacements:

  • JavaScript: encodeURIComponent(value) for values; encodeURI(url) for full URLs. Never the reverse.
  • Python: urllib.parse.quote(value, safe='') for values; urllib.parse.urlencode(params) builds whole query strings.
  • PHP: rawurlencode($value) (gives %20) for values in URLs; urlencode($value) (gives +) matches form encoding.
  • The trap: every one of these double-encodes if you pass them an already-encoded string. Encode once, at the point of assembly, and never again.

When NOT to Encode

Encoding has a mirror-image rule that causes just as many bugs: do not encode things that are already encoded, and do not encode values that were never meant for a URL.

  • Already-encoded strings: passing tea%20coffee through an encoder again produces tea%2520coffee. If a value arrives encoded, decode or leave it — never re-encode blindly.
  • Data going into a database or JSON payload: URL encoding is a transport format for URLs, not a general-purpose "make text safe" function. Encoding text before storing it just leaves %20 litter in your data.
  • Base64url tokens: JWTs and similar tokens already use an encoding designed to survive URLs; wrapping them in another layer of encoding breaks their signatures.

The pattern behind both mistakes is the same: encoding applied at the wrong layer. Encode once, at the boundary where a value enters a URL, and decode once, where it leaves.

The Five-Second Rule

If you remember nothing else: URLs can only safely carry A-Z, a-z, 0-9, and - _ . ~. Everything else inside a value gets percent-encoded — and you encode the values, never the URL's structure.

Do that consistently and you will never ship a broken share link, a corrupted redirect, or a mysteriously failing API call because of encoding again.

Why Encoding Exists: The Problem It Solves

URLs were designed for a small character set, but the web speaks every language and symbol. Encoding is the bridge: it represents arbitrary characters using only URL-safe ones, so a link containing spaces, accents, or symbols travels intact through every system that handles URLs — browsers, servers, emails, chats. Without it, links break silently: truncated at the first space, mangled in transit, or misinterpreted by the server. Every broken "weird character" link you have ever seen was an encoding failure.

The Anatomy of Percent-Encoding

The mechanism is simple: take the character’s bytes in UTF-8, and write each byte as a percent sign followed by two hexadecimal digits. A space becomes %20, an accented e becomes two byte-sequences, an emoji becomes four. Decoding reverses the process. The elegance is in the uniformity — one rule, every character, every language. The complexity is in knowing which characters need it and which must be left alone.

The Safe Set (Memorize This)

Unreserved characters — never encode: A-Z, a-z, 0-9, and - _ . ~ (hyphen, underscore, period, tilde). Everything else is either reserved (encode when it is data, leave when it is structure) or unsafe (always encode). This small safe set is the foundation of every encoding decision; when in doubt, encode — over-encoding a safe character is harmless, under-encoding an unsafe one breaks things.

Reserved Characters: Structure vs. Data

Characters like / ? # & = have jobs in URL structure: slashes separate paths, question marks start queries, ampersands separate parameters. When these characters appear as data (a search for "fish & chips"), they must be encoded (%26) so they are not mistaken for structure. The rule: structure stays literal, data gets encoded. Most encoding bugs are this rule violated — a literal & inside a parameter value, splitting one parameter into two.

Encoding in Practice: Common Scenarios

  • Query parameters: encode values, never the ? & = structure around them.
  • File names in URLs: spaces to %20 (or +, in form contexts — know the difference).
  • Non-English text: UTF-8 then percent-encode; modern browsers display it decoded, servers receive it encoded.
  • Redirect URLs as parameters: encode the entire inner URL — double-encoding territory, handle with care.
  • API calls: encode user input before interpolating into URLs; unencoded input is both a bug and a security risk.

The Classic Pitfalls

Double encoding: %20 encoded again becomes %2520 — the percent sign itself encoded. Symptom: literal "%20" visible in the page. Fix: encode exactly once, at the boundary. Encoding the structure: turning ? into %3F breaks the URL’s anatomy. Forgetting to encode: spaces and symbols truncated or mangled in transit. Mixing + and %20: + means space only in form-encoded query strings, not in paths — a classic source of subtle bugs.

Encoding and Security

Unencoded user input in URLs enables injection: crafted input can break out of parameters, alter queries, or smuggle content past naive validation. Proper encoding is a security boundary, not just hygiene — it ensures user data stays data and never becomes structure. Validate on the server regardless (encoding is not validation), but encode faithfully on the way out. The two practices together close a whole class of vulnerabilities.

Tools and Workflow

Do not hand-encode except to learn: use the URL encoder for one-off values, your language’s standard library functions in code (never string-replace hacks), and framework helpers that encode by default. The workflow rule: encode at the boundary where data enters the URL, decode where it leaves, and never in between. Centralize it — scattered ad-hoc encoding is where double-encoding breeds.

Debugging Encoding Issues

Symptoms and diagnoses: literal % sequences visible means double encoding or missing decode; truncated values mean unencoded special characters; "works in browser, fails in API" often means the browser was helpfully encoding what your code did not. Inspect the actual bytes on the wire (browser dev tools show them) rather than guessing. And reproduce with the simplest possible input first — complexity hides the single character causing the trouble.

Frequently Asked Questions

Should I encode the whole URL? No — encode the data parts, leave the structure (:// ? & = /) intact.

What is the difference between encodeURI and encodeURIComponent? In JavaScript: the first preserves structure (for whole URLs), the second encodes everything (for values). Using the wrong one is a classic bug.

Do modern browsers handle this automatically? Mostly for display and navigation — but programmatic URL construction in your code is still your responsibility.

Is + the same as %20? Only in form-encoded query strings. In paths, + is a literal plus. Know your context.

"Structure stays literal, data gets encoded, exactly once — the three-word religion of URLs that never breaks."

Your Next Step

Find a place in your project where URLs are built from user input — a search box, a share link, a redirect. Check whether the values are encoded at the boundary. If they are string-concatenated raw, you have found both a bug and a security gap, and the fix is one function call. The whole discipline of URL encoding fits in that single habit: encode data at the boundary, every time, and the weird-character link failures become someone else’s problem.

No ratings yet