Text processing failures share one property: they raise no error. String comparison returning false, a regex that never matches, CSV rows with misaligned columns — the program runs, the exit code is 0, and the problem only surfaces when the result is used. This article covers three frequent silent failure classes; tool support is listed in the in-site text processing reference.
The first class is invisible characters, which violate intuition because the character set you see is not the set that exists. Zero-width spaces (U+200B), zero-width joiners (U+200D), non-breaking spaces (U+00A0), byte order marks and curly quote variants all occupy no space and display nothing, yet they participate in every operation. So trim leaves content behind (U+00A0 is not whitespace as far as JavaScript's trim is concerned), ^abc$ fails because the string is really \u200Babc, two filenames look identical but differ, and a unique constraint collides for no visible reason.
The only reliable diagnostic is printing code points, never eyeballing. A text statistics tool lists them per character, making invisible differences obvious. The fix belongs at the boundary, not at each comparison site: strip zero-width characters and BOMs, convert U+00A0 to a plain space, unify quotes to ASCII, normalise to NFKC and unify line endings. NFKC deserves special mention because it folds fullwidth forms to ASCII and collapses compatibility characters — turning fullwidth Latin into plain letters and circled digits into plain digits — which matters greatly when handling text pasted from arbitrary sources.
The second class is CSV parsing. Splitting on newlines looks harmless and is wrong. RFC 4180 explicitly permits newlines inside fields when they are wrapped in double quotes, so a file containing a,b followed by "line1\nline2",c has three lines but only two records. Splitting on \n yields three fragments, the second being line2",c, and every downstream step now rests on a wrong record count. Worse, it raises nothing: split always succeeds and returns an array, so you believe you have the data.
Use a real parser instead — Python's csv module, papaparse, ExcelJS — which handle quote doubling, embedded newlines, delimiter escaping and optionally the BOM. Manual splitting is safe only when you know there are no quoted fields. Note also that splitting should be on the delimiter (comma, semicolon, tab), which is an operation orthogonal to line breaks; conflating the two is a common bug.
The third class is regex boundaries. Two properties are routinely misunderstood. First, a dot is not a literal dot: a.c matches abc, a1c and a-c, so using it to convert separators damages far more than intended. Second, global replacement has collateral effects. Consider "replace every dot with a comma" for European number formatting: decimal points change as intended, but version numbers become 1,2,3, domains become example,com and IP addresses become 192,168,1,1 — all certainly wrong.
The correct approach has three steps. Count before replacing: the length of str.match(re) reveals the blast radius, and a pattern expected to hit 5 sites returning 500 means the regex is wrong. Constrain with character classes: /\d+\.\d+/g matches only number-dot-number, leaving versions and IPs alone. Anchor with \b when matching whole words. Counting first takes seconds and catches nearly every malformed pattern, yet it is almost always skipped.
One more common trap: deduplicating by whole line. Given b,2 / a,1 / b,2, removing duplicate lines yields two rows — but if the duplication means the same user ID appeared twice, the real problem is not the repeated line but the data itself, which should be aggregated or rejected. Specify whether deduplication is by whole line or by a business key, because the two mean very different things.
A practical checklist: print code points when comparison fails rather than hunting by eye; normalise to NFKC and strip zero-width characters at ingestion; parse CSV with a real parser instead of splitting on newlines; count regex matches before replacing; constrain with classes like \d+\. instead of a bare dot; and decide explicitly whether deduplication is by line or by business key.