← Back to Blog

Text Processing: Invisible Characters, CSV Parsing and Regex Boundaries as Silent Failures

Text processing failures share one property: they raise no error. String comparison returning false, a regex that never matches, CSV rows with misaligned columns — the program runs, the exit code is 0, and the problem only surfaces when the result is used. This article covers three frequent silent failure classes; tool support is listed in the in-site text processing reference.

The first class is invisible characters, which violate intuition because the character set you see is not the set that exists. Zero-width spaces (U+200B), zero-width joiners (U+200D), non-breaking spaces (U+00A0), byte order marks and curly quote variants all occupy no space and display nothing, yet they participate in every operation. So trim leaves content behind (U+00A0 is not whitespace as far as JavaScript's trim is concerned), ^abc$ fails because the string is really \u200Babc, two filenames look identical but differ, and a unique constraint collides for no visible reason.

The only reliable diagnostic is printing code points, never eyeballing. A text statistics tool lists them per character, making invisible differences obvious. The fix belongs at the boundary, not at each comparison site: strip zero-width characters and BOMs, convert U+00A0 to a plain space, unify quotes to ASCII, normalise to NFKC and unify line endings. NFKC deserves special mention because it folds fullwidth forms to ASCII and collapses compatibility characters — turning fullwidth Latin into plain letters and circled digits into plain digits — which matters greatly when handling text pasted from arbitrary sources.

The second class is CSV parsing. Splitting on newlines looks harmless and is wrong. RFC 4180 explicitly permits newlines inside fields when they are wrapped in double quotes, so a file containing a,b followed by "line1\nline2",c has three lines but only two records. Splitting on \n yields three fragments, the second being line2",c, and every downstream step now rests on a wrong record count. Worse, it raises nothing: split always succeeds and returns an array, so you believe you have the data.

Use a real parser instead — Python's csv module, papaparse, ExcelJS — which handle quote doubling, embedded newlines, delimiter escaping and optionally the BOM. Manual splitting is safe only when you know there are no quoted fields. Note also that splitting should be on the delimiter (comma, semicolon, tab), which is an operation orthogonal to line breaks; conflating the two is a common bug.

The third class is regex boundaries. Two properties are routinely misunderstood. First, a dot is not a literal dot: a.c matches abc, a1c and a-c, so using it to convert separators damages far more than intended. Second, global replacement has collateral effects. Consider "replace every dot with a comma" for European number formatting: decimal points change as intended, but version numbers become 1,2,3, domains become example,com and IP addresses become 192,168,1,1 — all certainly wrong.

The correct approach has three steps. Count before replacing: the length of str.match(re) reveals the blast radius, and a pattern expected to hit 5 sites returning 500 means the regex is wrong. Constrain with character classes: /\d+\.\d+/g matches only number-dot-number, leaving versions and IPs alone. Anchor with \b when matching whole words. Counting first takes seconds and catches nearly every malformed pattern, yet it is almost always skipped.

One more common trap: deduplicating by whole line. Given b,2 / a,1 / b,2, removing duplicate lines yields two rows — but if the duplication means the same user ID appeared twice, the real problem is not the repeated line but the data itself, which should be aggregated or rejected. Specify whether deduplication is by whole line or by a business key, because the two mean very different things.

A practical checklist: print code points when comparison fails rather than hunting by eye; normalise to NFKC and strip zero-width characters at ingestion; parse CSV with a real parser instead of splitting on newlines; count regex matches before replacing; constrain with classes like \d+\. instead of a bare dot; and decide explicitly whether deduplication is by line or by business key.

Advertisement

Frequently Asked Questions

Why do two visually identical strings compare unequal?

**Almost always because of invisible characters.** Text copied from web pages, Word documents and PDFs often carries zero-width spaces (U+200B), zero-width joiners (U+200D), non-breaking spaces (U+00A0), byte order marks, and curly quote variants. These occupy no space and display nothing, yet they survive string comparison, regex matching and trim. Typical symptoms: content remaining after trim, `^abc$` failing to match, filenames that differ while looking identical. **Do not debug this by eye** — print the code points (a text statistics tool will show them per character). **Prevent it by normalising at the boundary**: strip zero-width characters, convert U+00A0 to a plain space, and apply NFKC.

Can I process CSV with split on newlines?

**Not unless you know the file has no quoted fields.** RFC 4180 explicitly permits newlines inside fields as long as they are wrapped in double quotes. So a line `a,b` followed by `"line1 line2",c` splits into two halves when you split on newlines, corrupting everything downstream — and silently. Use a real CSV parser (Python's `csv` module, papaparse, ExcelJS), which handle quote escaping, embedded newlines and delimiter escaping. Manual splitting is only safe for simple formats known to contain no quotes or newlines, or when you split on a delimiter such as comma, semicolon or tab instead of on line breaks.

What is wrong with using a dot to match a separator in a regex?

**A dot in a regex matches any character, not a literal dot.** `a.c` matches abc, a1c and a-c — anything three characters long starting with a and ending with c. Using it to convert separators therefore damages far more than intended. A subtler problem is **collateral damage from global replacement**: replacing every dot with a comma also hits decimal points, version numbers like 1.2.3, domain names and IP addresses. The correct approach: (1) count matches first via the length of `str.match(re)` to confirm the blast radius; (2) constrain boundaries with character classes, for example `/\d+\.\d+/g` matches only decimals and leaves version numbers alone; (3) anchor with `\b` where appropriate. A text statistics tool can count matches quickly, and this step should always come before replacing.

← Back to Blog