Encoding guide · Base64
Any bytes in.
Safe text out.
Not a secret.
Base64 turns arbitrary binary data — an image, a PDF, a cryptographic key — into plain ASCII text that survives systems built only for text: JSON, email, URLs, HTML attributes. It's everywhere, it's simple, and it's routinely mistaken for security. Here's exactly how it works. 👇
Text, images, files, keys — Base64 doesn't care what the bytes mean.
A–Z, a–z, 0–9, + and /, plus = for padding.
No key, no secret. Anyone can decode it in one line.
Base64 is encoding, not encryption. Base64-encoding a password, API key or personal data hides it from nobody. If you need confidentiality, use real encryption and access control — Base64 only changes how the bytes are written down.
02 · Under the hood
⚙️ How it works: 3 bytes → 4 characters
The idea is simple. Take the input three bytes (24 bits) at a time, split those 24 bits into four groups of 6 bits, and write each 6-bit number (0–63) as one character from a 64-character alphabet. Defined in RFC 4648, the standard alphabet is:
| Values | Characters |
|---|---|
| 0 – 25 | A B C … Z |
| 26 – 51 | a b c … z |
| 52 – 61 | 0 1 2 … 9 |
| 62 | + |
| 63 | / |
| padding | = |
🧮 Worked example: "Man" → "TWFu"
-
1️⃣ Text to bytes
In ASCII/UTF-8,
M= 77,a= 97,n= 110. -
2️⃣ Bytes to bits
010011010110000101101110— 24 bits in a row. -
3️⃣ Regroup into sixes
010011010110000101101110→ the numbers 19, 22, 5, 46. -
4️⃣ Look up the alphabet
19 =
T, 22 =W, 5 =F, 46 =u. Result: TWFu.
🧩 Where the = comes from
When the input length isn't a multiple of three, the last group is short. The encoder pads the missing bits with zeros and marks the gap with =:
| Input | Bytes | Output | Why |
|---|---|---|---|
| Man | 3 | TWFu | Full group — no padding |
| Ma | 2 | TWE= | 2 bytes → 3 characters + one = |
| M | 1 | TQ== | 1 byte → 2 characters + two == |
So a Base64 string never ends in more than two = signs, and a padded one always has a length that's a multiple of 4. A string that breaks either rule has been damaged or truncated.
03 · The cost
📏 The ~33% size overhead
Every 3 bytes become 4 characters, so encoded data is about four-thirds the size of the original. The exact padded length for n input bytes is 4 × ceil(n / 3).
| Original | Base64 |
|---|---|
| 3 bytes | 4 characters |
| 100 bytes | 136 characters |
| 30 KB image | 40 KB of text |
| 3 MB PDF | 4 MB of text |
MIME email adds a little more on top by wrapping encoded lines at 76 characters. HTTP compression claws some of the overhead back, but not all of it — and the decoded copy has to sit in memory alongside the encoded one.
Rule of thumb: Base64 is great for small binary values inside text formats. For real files, send the bytes directly — a multipart upload or a binary response body — rather than a giant string inside JSON.
04 · The important variant
🔗 Base64URL
Two standard characters cause trouble on the web: + is read as a space in query strings, and / is a path separator. = also has meaning in URLs. Base64URL (also in RFC 4648) swaps the two characters and usually drops the padding.
| What | Standard Base64 | Base64URL |
|---|---|---|
| Value 62 | + | - |
| Value 63 | / | _ |
| Padding | Usually kept | Usually omitted |
Example (Hi?>) | SGk/Pg== | SGk_Pg |
| Typical use | Email, data URIs, file payloads | JWTs, URL parameters, filenames |
The most common place you'll see it is JSON Web Tokens: each of the three dot-separated parts of a JWT is Base64URL. That means the header and claims of any JWT are readable by anyone who holds the token — paste one into the JWT Decoder and see. The signature protects against tampering, not reading.
Decoding fails with "invalid length"? You probably have unpadded Base64URL going into a strict standard decoder. Swap - → + and _ → /, then add = until the length is a multiple of 4.
05 · In the wild
🗺️ Where Base64 shows up
| Where | What it's doing |
|---|---|
| 🖼️ Data URIs | Embedding a small image straight into HTML or CSS: src="data:image/png;base64,iVBORw0KGgo…". Saves a request, costs ~33% extra bytes and can't be cached separately. |
| ✉️ Email attachments | MIME uses Base64 so binary attachments pass through mail systems designed for text. |
| 🔑 HTTP Basic auth | Authorization: Basic dXNlcjpwYXNz is just user:pass encoded — which is why Basic auth is only acceptable over HTTPS. |
| 🪪 Tokens & keys | JWTs, PEM certificates (the block between -----BEGIN lines), SSH public keys. |
| 📦 JSON payloads | Carrying small binary values — a signature, a thumbnail — where the format has no binary type. |
| ☸️ Config | Kubernetes Secrets store values Base64-encoded. That's encoding for transport, not protection. |
👍 Good fit
- Small binary values inside JSON, XML or HTML
- Tiny icons as data URIs
- Protocols that require it (MIME, Basic auth, JWT)
👎 Poor fit
- Large files — upload the bytes instead
- Hiding secrets — it hides nothing
- URL parameters — unless you use Base64URL
06 · The classic bug
🌏 Text, Unicode & the btoa trap
Base64 encodes bytes, not characters. Before you can encode text, you have to choose how to turn it into bytes — and the answer should almost always be UTF-8. The same text in a different encoding gives different Base64.
The browser's built-in btoa() predates that thinking. It accepts a string where every character must be in the Latin-1 range (code points 0–255) and treats each one as a byte. Anything else throws:
btoa("Hello") // "SGVsbG8=" fine btoa("नमस्ते") // ❌ InvalidCharacterError btoa("café") // "Y2Fm6Q==" — no error, but é was encoded as one Latin-1 byte, not UTF-8
The fix is to encode to UTF-8 bytes first, as in the JavaScript snippet in the next section. Encoded correctly, é is the two UTF-8 bytes C3 A9, which becomes w6k=.
Decoded data is untrusted. A Base64 string can hold anything — an executable, a script-laden SVG, a huge file. Check size and type before you save, render or run what comes out of a decoder.
07 · In your code
🛠️ Encode & decode snippets
// encode const bytes = new TextEncoder().encode("नमस्ते"); const encoded = btoa(String.fromCharCode(...bytes)); // decode const raw = Uint8Array.from(atob(encoded), c => c.charCodeAt(0)); const text = new TextDecoder().decode(raw); // newer engines also offer Uint8Array.prototype.toBase64() and // Uint8Array.fromBase64() — check browser support before relying on them
Buffer.from("Hello", "utf8").toString("base64"); // "SGVsbG8=" Buffer.from("Hello", "utf8").toString("base64url"); // "SGVsbG8" Buffer.from("SGVsbG8=", "base64").toString("utf8"); // "Hello"
using System.Text; string encoded = Convert.ToBase64String(Encoding.UTF8.GetBytes("Hello")); string decoded = Encoding.UTF8.GetString(Convert.FromBase64String(encoded)); // Base64URL: System.Buffers.Text.Base64Url (.NET 9+), or in ASP.NET Core // Microsoft.AspNetCore.WebUtilities.WebEncoders.Base64UrlEncode(bytes)
import base64 encoded = base64.b64encode("Hello".encode("utf-8")).decode("ascii") decoded = base64.b64decode(encoded).decode("utf-8") url_safe = base64.urlsafe_b64encode(data) # uses - and _ (keeps = padding)
# Linux / macOS — use printf or echo -n, or the trailing newline gets encoded too printf 'Hello' | base64 # SGVsbG8= echo 'SGVsbG8=' | base64 -d # Hello (older macOS: -D) base64 -w 0 photo.png > photo.b64 # GNU: no line wrapping # PowerShell [Convert]::ToBase64String([Text.Encoding]::UTF8.GetBytes("Hello"))
No code needed for a quick check. The Base64 Encoder / Decoder handles text and files in the browser. For values headed into a query string, you may want the URL Encoder / Decoder instead — or as well.