Text & encoding ยท String utilities explained
Small text jobs,
done right, without
writing code.
A string utility is anything that takes text in and gives cleaner, safer or differently shaped text back โ change the case, strip whitespace, make a URL slug, remove duplicate lines, encode it for a URL. Every developer does these jobs weekly. This page covers what the common operations are, when a browser tool beats a script, and the Unicode traps that catch both. ๐
Trim, collapse spaces, dedupe lines, remove emoji or invisible characters.
Case conversion, slugs, split and join, find and replace, sort lines.
URL, HTML, Base64, JSON and Unicode escapes for wherever the text is going next.
Who this is for: anyone who has pasted a column of names into a text editor and started fixing them by hand. If you then want the same logic in code, the C# string methods and JavaScript string methods guides map each operation to a real method.
02 ยท The vocabulary
๐งฐ The common string operations
Almost every text tool, and almost every string library, is a combination of the same dozen operations. Knowing their names makes it faster to find the right tool โ and the right method when you move to code.
| Operation | What it does | Typical use |
|---|---|---|
| ๐ Case conversion | Upper, lower, Title Case, Sentence case, and programmer cases: camelCase, PascalCase, snake_case, kebab-case, CONSTANT_CASE | Turning a spreadsheet header into a property name or an environment variable key. |
| โ๏ธ Trim & whitespace | Removes leading/trailing spaces, collapses runs of spaces, strips blank lines | Cleaning values copied from PDFs, emails and web pages. |
| ๐ Slug generation | Lowercases, removes accents and punctuation, joins words with hyphens | URL paths, file names, anchor ids. |
| ๐ช Split & join | Breaks text on a delimiter into a list, or glues a list back together | Comma list โ one-per-line; building an SQL IN (...) list. |
| ๐ Find & replace | Swaps one substring (or regex match) for another | Renaming a term everywhere; fixing a repeated typo. |
| ๐งฎ Dedupe & sort | Removes repeated lines, sorts alphabetically or naturally | Merging two email lists; cleaning log output. |
| ๐ Counting | Characters, words, lines, bytes | Meta description limits, tweet length, database column sizes. |
| ๐ Encoding & escaping | URL, HTML, Base64, JSON and Unicode escapes | Query strings, embedding text in HTML or JSON, API payloads. |
| โ๏ธ Compare | Shows what changed between two versions of a text | Config drift, "what did the client edit in this copy?" |
Order matters when you chain them. Trim before you dedupe (otherwise "apple" and "apple " both survive), and normalise case before you sort if you want Apple and apple next to each other.
03 ยท Before and after
๐ What each operation actually produces
Concrete input and output beats a definition. Every row below is something you can reproduce in the String Tools page or its siblings.
| Operation | Input | Output |
|---|---|---|
| snake_case | Customer Order Number | customer_order_number |
| camelCase | Customer Order Number | customerOrderNumber |
| PascalCase | customer order number | CustomerOrderNumber |
| CONSTANT_CASE | api base url | API_BASE_URL |
| kebab-case | Primary Button Color | primary-button-color |
| Slug | Crรจme Brรปlรฉe: 10 Tips! | creme-brulee-10-tips |
| Trim + collapse | " John Smith " | "John Smith" |
| Split (comma โ lines) | red,green,blue | red green blue |
| Dedupe lines | a b a c | a b c |
| URL encode | name=Ana Marรญa&x=1 | name%3DAna%20Mar%C3%ADa%26x%3D1 |
| HTML encode | <b>"Hi"</b> | <b>"Hi"</b> |
| Base64 | hello | aGVsbG8= |
| Unicode escape | cafรฉ | cafรฉ |
๐ Encoding is not encryption
URL encoding, HTML encoding and Base64 are all reversible by anyone. They exist so text survives a trip through a system that would otherwise misread it โ a space in a query string, a < in HTML, binary bytes in a JSON field. None of them hide anything. If you need secrecy, you need encryption or hashing, not an encoder.
๐ Pick the encoding for the destination
- Going into a URL? URL (percent) encoding.
- Going into HTML? HTML entity encoding.
- Going into JSON? JSON string escaping (
\",\n). - Binary in a text field? Base64.
๐ซ Common mix-ups
- Encoding twice โ
%20becomes%2520. - Using HTML encoding for a URL parameter.
- Treating Base64 as a password store.
- Form-style
+for spaces where the server expects%20.
04 ยท Decide per job
โ๏ธ A browser tool or a line of code?
Both are the right answer โ for different jobs. The deciding question is "will this happen again, automatically?"
๐งฐ Use a tool whenโฆ
- It's a one-off: cleaning a list someone emailed you.
- You want to see the result before trusting it.
- You're checking what your code should output.
- The person doing it isn't a developer.
- You're debugging an encoded value from a log or a URL.
๐ป Write code whenโฆ
- It runs on every request, import or build.
- The input is large or arrives continuously.
- Rules are specific to your domain (your slug format, your casing).
- It has to be tested and reviewed.
- The data is sensitive enough that it shouldn't leave your system at all.
The best workflow is often both: prototype the transformation in a tool until the output looks right, then write the code and use the tool's output as the expected value in a unit test.
Be careful what you paste anywhere. Tokens, connection strings and customer data should not go into any website you haven't vetted. QuickDeveloperTools' string tools run in your browser, but the safe habit is the same everywhere: redact secrets before pasting.
05 ยท Where it goes wrong
๐ Unicode pitfalls every string tool hits
Most string bugs are not logic bugs. They're a mismatch between what a person calls one character and what the computer counts. There are three layers, and they don't line up:
| Layer | What it is | Example |
|---|---|---|
| ๐งฑ Code unit | What .length counts in JavaScript and C# โ a 16-bit UTF-16 unit | "๐".length is 2 |
| ๐ข Code point | One Unicode number, e.g. U+1F600 | ๐ is 1 code point |
| ๐๏ธ Grapheme cluster | What a reader sees as one character | ๐จโ๐ฉโ๐ง is 1 grapheme, 5 code points, 8 UTF-16 units |
1๏ธโฃ Surrogate pairs
UTF-16 can only fit code points up to U+FFFF in one unit. Everything above โ most emoji, many CJK extension characters, mathematical letters โ is stored as a surrogate pair of two units. Cut a string in the middle of a pair (a naive substring, a character-by-character reverse, truncation to "140 characters") and you get a broken ๏ฟฝ character.
2๏ธโฃ Normalization
The letter รฉ can be stored as one code point (U+00E9) or as e plus a combining acute accent (U+0065 U+0301). They look identical and compare unequal. Text pasted from macOS file names, some PDFs and some keyboards arrives in the decomposed form.
// JavaScript "cafรฉ" === "cafeฬ" // false "cafรฉ".normalize() === "cafeฬ".normalize() // true (both NFC) // C# "cafรฉ" == "cafeฬ" // false "cafeฬ".Normalize() == "cafรฉ" // true
Rule of thumb: normalize to NFC before you compare, dedupe or store user-entered text. Slug generators use the opposite trick โ decompose to NFD, then drop the combining marks โ which is how Crรจme becomes creme.
3๏ธโฃ Grapheme clusters
Skin-tone emoji (๐๐ฝ = thumbs up + modifier), flags (๐ฎ๐ณ = two regional indicator letters) and family emoji joined with zero-width joiners are several code points that render as one symbol. If you're enforcing a visible character limit, count graphemes, not .length:
// JavaScript โ Intl.Segmenter (all current browsers and Node) [...new Intl.Segmenter().segment("๐จโ๐ฉโ๐ง hi")].length // 4 // C# โ StringInfo is grapheme-aware since .NET 5 new StringInfo("๐จโ๐ฉโ๐ง hi").LengthInTextElements // 4
๐ซฅ Invisible characters
| Character | Code point | Why it bites |
|---|---|---|
| Non-breaking space | U+00A0 | Looks like a space; breaks split(" ") and exact matches. Common in text copied from web pages and Word. |
| Zero-width space | U+200B | Completely invisible, and not treated as whitespace by trim() in JS or Trim() in .NET. |
| Byte order mark | U+FEFF | Sneaks onto the start of files saved as "UTF-8 with BOM"; the first CSV header then fails to match. |
| Zero-width joiner | U+200D | Legitimately glues emoji together โ stripping it splits ๐จโ๐ฉโ๐ง into three people. |
The Unicode Converter shows the code points in a string, which is the fastest way to find out why two "identical" values don't match.
Case conversion is language-dependent too. In Turkish, the uppercase of i is ฤฐ, not I. In JavaScript, "ร".toUpperCase() returns "SS" โ a longer string. For identifiers, keys and protocol values, always use the invariant or ordinal form rather than the user's culture.
06 ยท From tool to code
๐งโ๐ป The same operations in C# and JavaScript
Once a transformation looks right, this is roughly what it becomes in code. The two language guides go much deeper, including the gotchas.
| Operation | C# | JavaScript |
|---|---|---|
| Upper / lower | s.ToUpperInvariant() | s.toUpperCase() |
| Trim | s.Trim() | s.trim() |
| Replace all | s.Replace("a", "b") | s.replaceAll("a", "b") |
| Split | s.Split(',') | s.split(",") |
| Join | string.Join(",", list) | list.join(",") |
| Contains (ignore case) | s.Contains("x", StringComparison.OrdinalIgnoreCase) | s.toLowerCase().includes("x") |
| Dedupe lines | lines.Distinct() | [...new Set(lines)] |
| Normalize | s.Normalize() | s.normalize() |
| URL encode | Uri.EscapeDataString(s) | encodeURIComponent(s) |
| Base64 | Convert.ToBase64String(Encoding.UTF8.GetBytes(s)) | btoa(s) (Latin-1 only) |
๐ค C# string methods
Inspect, search, compare with StringComparison, transform, split and join, formatting, StringBuilder and spans.
๐จ JavaScript string methods
slice vs substring, replace vs replaceAll, padding, localeCompare, template literals and emoji-safe iteration.
About that btoa row: btoa throws on any character above U+00FF, so btoa("cafรฉ โ") fails. Encode to UTF-8 bytes first (TextEncoder) โ or use the Base64 Encoder / Decoder to check the expected output.
07 ยท On this site
๐งฐ The string toolbox on QuickDeveloperTools
Each tool does one job and shows the result instantly, so you can check the output before you copy it anywhere.
| Tool | Reach for it when |
|---|---|
| String Tools | Case conversion, trimming, line tools, reversing and text statistics in one place โ the general-purpose starting point. |
| Slug Generator | You need a URL path, file name or anchor id from a title. |
| Find & Replace | Bulk replacements, with or without regular expressions. |
| Remove Duplicates | A list has repeated lines and you want each value once. |
| Duplicate Word Finder | Proofreading for "the the" and overused words. |
| Unicode Converter | Seeing code points and escapes โ finding that hidden U+200B. |
| Emoji Remover | A system can't store emoji, or you need plain text for a legacy field. |
| Regex Tester | Building the pattern for a find-and-replace or a validation rule. |
| Text Compare | Two versions of a text and you need to see exactly what changed. |
| URL Encoder / Decoder | Building or debugging query strings. |
| HTML Encoder / Decoder | Showing markup as text, or reading entity-encoded content. |
| Base64 Encoder / Decoder | Inspecting a Base64 payload or preparing one. |
If you remember one thing: trim, normalize and pick an explicit comparison before doing anything clever. Most "the strings are equal but the code says they're not" bugs disappear after those three steps.