Skip to content
Browse tools

Encoding guide · Base64

Any bytes in.
Safe text out.
Not a secret.

Base64 turns arbitrary binary data — an image, a PDF, a cryptographic key — into plain ASCII text that survives systems built only for text: JSON, email, URLs, HTML attributes. It's everywhere, it's simple, and it's routinely mistaken for security. Here's exactly how it works. 👇

📥
Input
Any bytes

Text, images, files, keys — Base64 doesn't care what the bytes mean.

🔤
Output
64 safe characters

A–Z, a–z, 0–9, + and /, plus = for padding.

🔓
Security
None at all

No key, no secret. Anyone can decode it in one line.

🚨

Base64 is encoding, not encryption. Base64-encoding a password, API key or personal data hides it from nobody. If you need confidentiality, use real encryption and access control — Base64 only changes how the bytes are written down.

02 · Under the hood

⚙️ How it works: 3 bytes → 4 characters

The idea is simple. Take the input three bytes (24 bits) at a time, split those 24 bits into four groups of 6 bits, and write each 6-bit number (0–63) as one character from a 64-character alphabet. Defined in RFC 4648, the standard alphabet is:

ValuesCharacters
0 – 25A B C … Z
26 – 51a b c … z
52 – 610 1 2 … 9
62+
63/
padding=

🧮 Worked example: "Man" → "TWFu"

  1. 1️⃣ Text to bytes

    In ASCII/UTF-8, M = 77, a = 97, n = 110.

  2. 2️⃣ Bytes to bits

    01001101 01100001 01101110 — 24 bits in a row.

  3. 3️⃣ Regroup into sixes

    010011 010110 000101 101110 → the numbers 19, 22, 5, 46.

  4. 4️⃣ Look up the alphabet

    19 = T, 22 = W, 5 = F, 46 = u. Result: TWFu.

🧩 Where the = comes from

When the input length isn't a multiple of three, the last group is short. The encoder pads the missing bits with zeros and marks the gap with =:

InputBytesOutputWhy
Man3TWFuFull group — no padding
Ma2TWE=2 bytes → 3 characters + one =
M1TQ==1 byte → 2 characters + two ==

So a Base64 string never ends in more than two = signs, and a padded one always has a length that's a multiple of 4. A string that breaks either rule has been damaged or truncated.

03 · The cost

📏 The ~33% size overhead

Every 3 bytes become 4 characters, so encoded data is about four-thirds the size of the original. The exact padded length for n input bytes is 4 × ceil(n / 3).

OriginalBase64
3 bytes4 characters
100 bytes136 characters
30 KB image40 KB of text
3 MB PDF4 MB of text

MIME email adds a little more on top by wrapping encoded lines at 76 characters. HTTP compression claws some of the overhead back, but not all of it — and the decoded copy has to sit in memory alongside the encoded one.

💡

Rule of thumb: Base64 is great for small binary values inside text formats. For real files, send the bytes directly — a multipart upload or a binary response body — rather than a giant string inside JSON.

04 · The important variant

🔗 Base64URL

Two standard characters cause trouble on the web: + is read as a space in query strings, and / is a path separator. = also has meaning in URLs. Base64URL (also in RFC 4648) swaps the two characters and usually drops the padding.

WhatStandard Base64Base64URL
Value 62+-
Value 63/_
PaddingUsually keptUsually omitted
Example (Hi?>)SGk/Pg==SGk_Pg
Typical useEmail, data URIs, file payloadsJWTs, URL parameters, filenames

The most common place you'll see it is JSON Web Tokens: each of the three dot-separated parts of a JWT is Base64URL. That means the header and claims of any JWT are readable by anyone who holds the token — paste one into the JWT Decoder and see. The signature protects against tampering, not reading.

🧷

Decoding fails with "invalid length"? You probably have unpadded Base64URL going into a strict standard decoder. Swap - → + and _ → /, then add = until the length is a multiple of 4.

05 · In the wild

🗺️ Where Base64 shows up

WhereWhat it's doing
🖼️ Data URIsEmbedding a small image straight into HTML or CSS: src="data:image/png;base64,iVBORw0KGgo…". Saves a request, costs ~33% extra bytes and can't be cached separately.
✉️ Email attachmentsMIME uses Base64 so binary attachments pass through mail systems designed for text.
🔑 HTTP Basic authAuthorization: Basic dXNlcjpwYXNz is just user:pass encoded — which is why Basic auth is only acceptable over HTTPS.
🪪 Tokens & keysJWTs, PEM certificates (the block between -----BEGIN lines), SSH public keys.
📦 JSON payloadsCarrying small binary values — a signature, a thumbnail — where the format has no binary type.
☸️ ConfigKubernetes Secrets store values Base64-encoded. That's encoding for transport, not protection.

👍 Good fit

  • Small binary values inside JSON, XML or HTML
  • Tiny icons as data URIs
  • Protocols that require it (MIME, Basic auth, JWT)

👎 Poor fit

  • Large files — upload the bytes instead
  • Hiding secrets — it hides nothing
  • URL parameters — unless you use Base64URL

06 · The classic bug

🌏 Text, Unicode & the btoa trap

Base64 encodes bytes, not characters. Before you can encode text, you have to choose how to turn it into bytes — and the answer should almost always be UTF-8. The same text in a different encoding gives different Base64.

The browser's built-in btoa() predates that thinking. It accepts a string where every character must be in the Latin-1 range (code points 0–255) and treats each one as a byte. Anything else throws:

▸ browser console
btoa("Hello")     // "SGVsbG8="  fine
btoa("नमस्ते")    // ❌ InvalidCharacterError
btoa("café")      // "Y2Fm6Q==" — no error, but é was encoded as one Latin-1 byte, not UTF-8

The fix is to encode to UTF-8 bytes first, as in the JavaScript snippet in the next section. Encoded correctly, é is the two UTF-8 bytes C3 A9, which becomes w6k=.

🧨

Decoded data is untrusted. A Base64 string can hold anything — an executable, a script-laden SVG, a huge file. Check size and type before you save, render or run what comes out of a decoder.

07 · In your code

🛠️ Encode & decode snippets

▸ JavaScript · browser (UTF-8 safe)
// encode
const bytes   = new TextEncoder().encode("नमस्ते");
const encoded = btoa(String.fromCharCode(...bytes));

// decode
const raw  = Uint8Array.from(atob(encoded), c => c.charCodeAt(0));
const text = new TextDecoder().decode(raw);

// newer engines also offer Uint8Array.prototype.toBase64() and
// Uint8Array.fromBase64() — check browser support before relying on them
▸ Node.js
Buffer.from("Hello", "utf8").toString("base64");     // "SGVsbG8="
Buffer.from("Hello", "utf8").toString("base64url");  // "SGVsbG8"
Buffer.from("SGVsbG8=", "base64").toString("utf8");  // "Hello"
▸ C# / .NET
using System.Text;

string encoded = Convert.ToBase64String(Encoding.UTF8.GetBytes("Hello"));
string decoded = Encoding.UTF8.GetString(Convert.FromBase64String(encoded));

// Base64URL: System.Buffers.Text.Base64Url (.NET 9+), or in ASP.NET Core
// Microsoft.AspNetCore.WebUtilities.WebEncoders.Base64UrlEncode(bytes)
▸ Python
import base64

encoded = base64.b64encode("Hello".encode("utf-8")).decode("ascii")
decoded = base64.b64decode(encoded).decode("utf-8")

url_safe = base64.urlsafe_b64encode(data)   # uses - and _ (keeps = padding)
▸ terminal
# Linux / macOS — use printf or echo -n, or the trailing newline gets encoded too
printf 'Hello' | base64            # SGVsbG8=
echo 'SGVsbG8=' | base64 -d        # Hello  (older macOS: -D)
base64 -w 0 photo.png > photo.b64   # GNU: no line wrapping

# PowerShell
[Convert]::ToBase64String([Text.Encoding]::UTF8.GetBytes("Hello"))
⚡

No code needed for a quick check. The Base64 Encoder / Decoder handles text and files in the browser. For values headed into a query string, you may want the URL Encoder / Decoder instead — or as well.

📌 Alphabets, padding rules and the Base64URL variant follow RFC 4648. Library APIs reflect current JavaScript, Node.js, .NET and Python releases.

What is Base64?
7 min read