Skip to content
Browse tools

Text & encoding ยท String utilities explained

Small text jobs,
done right, without
writing code.

A string utility is anything that takes text in and gives cleaner, safer or differently shaped text back โ€” change the case, strip whitespace, make a URL slug, remove duplicate lines, encode it for a URL. Every developer does these jobs weekly. This page covers what the common operations are, when a browser tool beats a script, and the Unicode traps that catch both. ๐Ÿ‘‡

๐Ÿงน
Clean
Tidy text

Trim, collapse spaces, dedupe lines, remove emoji or invisible characters.

๐Ÿ”
Reshape
Change form

Case conversion, slugs, split and join, find and replace, sort lines.

๐Ÿ”
Encode
Make it safe

URL, HTML, Base64, JSON and Unicode escapes for wherever the text is going next.

๐Ÿ‘‹

Who this is for: anyone who has pasted a column of names into a text editor and started fixing them by hand. If you then want the same logic in code, the C# string methods and JavaScript string methods guides map each operation to a real method.

02 ยท The vocabulary

๐Ÿงฐ The common string operations

Almost every text tool, and almost every string library, is a combination of the same dozen operations. Knowing their names makes it faster to find the right tool โ€” and the right method when you move to code.

OperationWhat it doesTypical use
๐Ÿ”  Case conversionUpper, lower, Title Case, Sentence case, and programmer cases: camelCase, PascalCase, snake_case, kebab-case, CONSTANT_CASETurning a spreadsheet header into a property name or an environment variable key.
โœ‚๏ธ Trim & whitespaceRemoves leading/trailing spaces, collapses runs of spaces, strips blank linesCleaning values copied from PDFs, emails and web pages.
๐Ÿ”— Slug generationLowercases, removes accents and punctuation, joins words with hyphensURL paths, file names, anchor ids.
๐Ÿช“ Split & joinBreaks text on a delimiter into a list, or glues a list back togetherComma list โ†” one-per-line; building an SQL IN (...) list.
๐Ÿ” Find & replaceSwaps one substring (or regex match) for anotherRenaming a term everywhere; fixing a repeated typo.
๐Ÿงฎ Dedupe & sortRemoves repeated lines, sorts alphabetically or naturallyMerging two email lists; cleaning log output.
๐Ÿ“ CountingCharacters, words, lines, bytesMeta description limits, tweet length, database column sizes.
๐Ÿ” Encoding & escapingURL, HTML, Base64, JSON and Unicode escapesQuery strings, embedding text in HTML or JSON, API payloads.
โ†”๏ธ CompareShows what changed between two versions of a textConfig drift, "what did the client edit in this copy?"
๐Ÿ’ก

Order matters when you chain them. Trim before you dedupe (otherwise "apple" and "apple " both survive), and normalise case before you sort if you want Apple and apple next to each other.

03 ยท Before and after

๐Ÿ” What each operation actually produces

Concrete input and output beats a definition. Every row below is something you can reproduce in the String Tools page or its siblings.

OperationInputOutput
snake_caseCustomer Order Numbercustomer_order_number
camelCaseCustomer Order NumbercustomerOrderNumber
PascalCasecustomer order numberCustomerOrderNumber
CONSTANT_CASEapi base urlAPI_BASE_URL
kebab-casePrimary Button Colorprimary-button-color
SlugCrรจme Brรปlรฉe: 10 Tips!creme-brulee-10-tips
Trim + collapse"  John   Smith ""John Smith"
Split (comma โ†’ lines)red,green,bluered
green
blue
Dedupe linesa
b
a
c
a
b
c
URL encodename=Ana Marรญa&x=1name%3DAna%20Mar%C3%ADa%26x%3D1
HTML encode<b>"Hi"</b>&lt;b&gt;&quot;Hi&quot;&lt;/b&gt;
Base64helloaGVsbG8=
Unicode escapecafรฉcafรฉ

๐Ÿ” Encoding is not encryption

URL encoding, HTML encoding and Base64 are all reversible by anyone. They exist so text survives a trip through a system that would otherwise misread it โ€” a space in a query string, a < in HTML, binary bytes in a JSON field. None of them hide anything. If you need secrecy, you need encryption or hashing, not an encoder.

๐ŸŒ Pick the encoding for the destination

  • Going into a URL? URL (percent) encoding.
  • Going into HTML? HTML entity encoding.
  • Going into JSON? JSON string escaping (\", \n).
  • Binary in a text field? Base64.

๐Ÿšซ Common mix-ups

  • Encoding twice โ€” %20 becomes %2520.
  • Using HTML encoding for a URL parameter.
  • Treating Base64 as a password store.
  • Form-style + for spaces where the server expects %20.

04 ยท Decide per job

โš–๏ธ A browser tool or a line of code?

Both are the right answer โ€” for different jobs. The deciding question is "will this happen again, automatically?"

๐Ÿงฐ Use a tool whenโ€ฆ

  • It's a one-off: cleaning a list someone emailed you.
  • You want to see the result before trusting it.
  • You're checking what your code should output.
  • The person doing it isn't a developer.
  • You're debugging an encoded value from a log or a URL.

๐Ÿ’ป Write code whenโ€ฆ

  • It runs on every request, import or build.
  • The input is large or arrives continuously.
  • Rules are specific to your domain (your slug format, your casing).
  • It has to be tested and reviewed.
  • The data is sensitive enough that it shouldn't leave your system at all.

The best workflow is often both: prototype the transformation in a tool until the output looks right, then write the code and use the tool's output as the expected value in a unit test.

๐Ÿ”’

Be careful what you paste anywhere. Tokens, connection strings and customer data should not go into any website you haven't vetted. QuickDeveloperTools' string tools run in your browser, but the safe habit is the same everywhere: redact secrets before pasting.

05 ยท Where it goes wrong

๐ŸŒ Unicode pitfalls every string tool hits

Most string bugs are not logic bugs. They're a mismatch between what a person calls one character and what the computer counts. There are three layers, and they don't line up:

LayerWhat it isExample
๐Ÿงฑ Code unitWhat .length counts in JavaScript and C# โ€” a 16-bit UTF-16 unit"๐Ÿ˜€".length is 2
๐Ÿ”ข Code pointOne Unicode number, e.g. U+1F600๐Ÿ˜€ is 1 code point
๐Ÿ‘๏ธ Grapheme clusterWhat a reader sees as one character๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง is 1 grapheme, 5 code points, 8 UTF-16 units

1๏ธโƒฃ Surrogate pairs

UTF-16 can only fit code points up to U+FFFF in one unit. Everything above โ€” most emoji, many CJK extension characters, mathematical letters โ€” is stored as a surrogate pair of two units. Cut a string in the middle of a pair (a naive substring, a character-by-character reverse, truncation to "140 characters") and you get a broken ๏ฟฝ character.

2๏ธโƒฃ Normalization

The letter รฉ can be stored as one code point (U+00E9) or as e plus a combining acute accent (U+0065 U+0301). They look identical and compare unequal. Text pasted from macOS file names, some PDFs and some keyboards arrives in the decomposed form.

โ–ธ the same word, twice
// JavaScript
"cafรฉ" === "cafeฬ"                           // false
"cafรฉ".normalize() === "cafeฬ".normalize()   // true (both NFC)

// C#
"cafรฉ" == "cafeฬ"                            // false
"cafeฬ".Normalize() == "cafรฉ"                 // true

Rule of thumb: normalize to NFC before you compare, dedupe or store user-entered text. Slug generators use the opposite trick โ€” decompose to NFD, then drop the combining marks โ€” which is how Crรจme becomes creme.

3๏ธโƒฃ Grapheme clusters

Skin-tone emoji (๐Ÿ‘๐Ÿฝ = thumbs up + modifier), flags (๐Ÿ‡ฎ๐Ÿ‡ณ = two regional indicator letters) and family emoji joined with zero-width joiners are several code points that render as one symbol. If you're enforcing a visible character limit, count graphemes, not .length:

โ–ธ counting what the user sees
// JavaScript โ€” Intl.Segmenter (all current browsers and Node)
[...new Intl.Segmenter().segment("๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง hi")].length   // 4

// C# โ€” StringInfo is grapheme-aware since .NET 5
new StringInfo("๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง hi").LengthInTextElements          // 4

๐Ÿซฅ Invisible characters

CharacterCode pointWhy it bites
Non-breaking spaceU+00A0Looks like a space; breaks split(" ") and exact matches. Common in text copied from web pages and Word.
Zero-width spaceU+200BCompletely invisible, and not treated as whitespace by trim() in JS or Trim() in .NET.
Byte order markU+FEFFSneaks onto the start of files saved as "UTF-8 with BOM"; the first CSV header then fails to match.
Zero-width joinerU+200DLegitimately glues emoji together โ€” stripping it splits ๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง into three people.

The Unicode Converter shows the code points in a string, which is the fastest way to find out why two "identical" values don't match.

๐ŸŒ

Case conversion is language-dependent too. In Turkish, the uppercase of i is ฤฐ, not I. In JavaScript, "รŸ".toUpperCase() returns "SS" โ€” a longer string. For identifiers, keys and protocol values, always use the invariant or ordinal form rather than the user's culture.

06 ยท From tool to code

๐Ÿง‘โ€๐Ÿ’ป The same operations in C# and JavaScript

Once a transformation looks right, this is roughly what it becomes in code. The two language guides go much deeper, including the gotchas.

OperationC#JavaScript
Upper / lowers.ToUpperInvariant()s.toUpperCase()
Trims.Trim()s.trim()
Replace alls.Replace("a", "b")s.replaceAll("a", "b")
Splits.Split(',')s.split(",")
Joinstring.Join(",", list)list.join(",")
Contains (ignore case)s.Contains("x", StringComparison.OrdinalIgnoreCase)s.toLowerCase().includes("x")
Dedupe lineslines.Distinct()[...new Set(lines)]
Normalizes.Normalize()s.normalize()
URL encodeUri.EscapeDataString(s)encodeURIComponent(s)
Base64Convert.ToBase64String(Encoding.UTF8.GetBytes(s))btoa(s) (Latin-1 only)

๐Ÿ”ค C# string methods

Inspect, search, compare with StringComparison, transform, split and join, formatting, StringBuilder and spans.

Read the C# guide โ†’

๐ŸŸจ JavaScript string methods

slice vs substring, replace vs replaceAll, padding, localeCompare, template literals and emoji-safe iteration.

Read the JavaScript guide โ†’

โš ๏ธ

About that btoa row: btoa throws on any character above U+00FF, so btoa("cafรฉ โ˜•") fails. Encode to UTF-8 bytes first (TextEncoder) โ€” or use the Base64 Encoder / Decoder to check the expected output.

07 ยท On this site

๐Ÿงฐ The string toolbox on QuickDeveloperTools

Each tool does one job and shows the result instantly, so you can check the output before you copy it anywhere.

ToolReach for it when
String ToolsCase conversion, trimming, line tools, reversing and text statistics in one place โ€” the general-purpose starting point.
Slug GeneratorYou need a URL path, file name or anchor id from a title.
Find & ReplaceBulk replacements, with or without regular expressions.
Remove DuplicatesA list has repeated lines and you want each value once.
Duplicate Word FinderProofreading for "the the" and overused words.
Unicode ConverterSeeing code points and escapes โ€” finding that hidden U+200B.
Emoji RemoverA system can't store emoji, or you need plain text for a legacy field.
Regex TesterBuilding the pattern for a find-and-replace or a validation rule.
Text CompareTwo versions of a text and you need to see exactly what changed.
URL Encoder / DecoderBuilding or debugging query strings.
HTML Encoder / DecoderShowing markup as text, or reading entity-encoded content.
Base64 Encoder / DecoderInspecting a Base64 payload or preparing one.
๐Ÿ

If you remember one thing: trim, normalize and pick an explicit comparison before doing anything clever. Most "the strings are equal but the code says they're not" bugs disappear after those three steps.

๐Ÿ“Œ Unicode behaviour described here follows the Unicode Standard and current JavaScript and .NET runtimes. Library defaults can differ โ€” check the output of your own code with a known tricky input (an emoji, an accented letter, a non-breaking space) before relying on it.

String utilities
9 min read