← Back to Blog

Encryption vs Hashing vs Encoding: One Table to Tell Them Apart

One Table to Tell Them Apart

Aspect Encoding Hashing Encryption
Purpose Fit a channel/format Verify integrity / fingerprint Confidentiality
Reversibility Reversible (public table) Irreversible (information discarded) Reversible (needs a key)
Key None None (password hashing adds a salt) Required, kept secret
Output Text within a charset Fixed-length digest Ciphertext the same size as plaintext
Same input Deterministic Deterministic Deterministic modes yes; random IV varies
Examples Base64, URL encoding, Hex MD5, SHA-256, bcrypt AES-GCM, ChaCha20, RSA
Typical use Transport compatibility File verification, passwords Data confidentiality

One line: encoding re-represents, hashing fingerprints, encryption locks. Locks need keys, fingerprints cannot be reversed, and anyone can re-represent.


Encoding Is Not Encryption: Provable in 30 Seconds

「Base64 output looks like gibberish, so it must be encrypted」is the most common confusion. The gibberish feeling comes only from a different alphabet — the table is a public standard and decoding is one step:

echo -n "admin:123456" | base64     # → YWRtaW46MTIzNDU2
echo "YWRtaW46MTIzNDU2" | base64 -d # → admin:123456 (anyone, one step)

No key involved, so no confidentiality. Base64's proper role is transport compatibility: letting binary data pass through text-only channels (email bodies, JSON strings). See Base64 Encoding Explained.

URL encoding (%E4%BD%A0) and Unicode escapes (\u4f60) are the same category — they solve "can this character legally appear in this format", not secrecy.


Why Hashing Is Irreversible: Information Is Discarded

Another confusion: 「MD5 is irreversible, so what are those MD5 decryption sites?」

SHA-256 compresses input of any length into a fixed 32-byte digest. Along the way information is discarded: the 256-bit output space has 2²⁵⁶ possibilities while inputs are infinite — infinitely many different inputs can share one digest. There simply is not enough information in a digest to reconstruct the original.

"Decryption sites" therefore do not reverse anything — they look up: they precompute hashes of massive candidate string lists and match yours. That is the entire reason bare MD5/SHA-256 password storage is unsafe.

The right posture for passwords: salt + slow hash (bcrypt / argon2 / scrypt). Salt makes the same password hash differently per user (breaking rainbow tables), and the slow hash raises per-guess cost from nanoseconds to tens of milliseconds. Comparison and selection in the password hashing guide.


The Core of Encryption Is the Key, Not the Algorithm

Encryption algorithms are all public (the AES and ChaCha20 specs are one search away), and confidentiality depends entirely on the key. Three often-ignored corollaries:

  1. Key management fails more often than algorithm choice: keys hardcoded in source, committed in config files — more common and more fatal than "used ECB mode".
  2. Public algorithms are not a weakness: Kerckhoffs' principle — a system should remain secure when everything but the key is public.
  3. Mode choice matters equally: AES-ECB produces identical ciphertext blocks for identical plaintext blocks (the famous "ECB penguin"); AES-256-GCM gives you both confidentiality and integrity. Mode comparison in AES encryption modes.

How to Choose: Three Scenarios

Task Right tool Not
Store login passwords bcrypt / argon2 (salted slow hash) MD5, Base64, AES (reversible = retrievable)
Verify a downloaded file SHA-256 against the published digest MD5 (collisions are constructible)
Push binary through a text channel Base64 / Base64URL encryption (don't add key management to a keyless job)
Sensitive fields in a database AES-GCM authenticated encryption Base64, hashing (irreversible means you can never read it back)
Tamper-proof API payloads HMAC-SHA256 (keyed hash) concatenated bare MD5

Row three trips people in the other direction: a hash's irreversibility means data you "stored" as a hash can never be read back — hashing a phone number for later display is impossible by design.


Reproducible Results

The same input hello world through all three, verified in a terminal in 30 seconds:

import base64, hashlib
msg = b'hello world'

print(base64.b64encode(msg).decode())      # aGVsbG8gd29ybGQ=   ← encoding: reversible, no key
print(base64.b64decode(base64.b64encode(msg)))  # b'hello world' ← one step back
print(hashlib.sha256(msg).hexdigest())     # b94d27b9934d3e08…  ← hashing: irreversible
print(len(hashlib.sha256(msg).digest()))   # 32 (256 bit)

Then a real AES-CBC encryption with OpenSSL (key and IV fixed for reproducibility):

echo -n "hello world" | openssl enc -aes-128-cbc \
  -K 000102030405060708090a0b0c0d0e0f \
  -iv f0f1f2f3f4f5f6f7f8f9fafbfcfdfeff | xxd -p
# → ciphertext; changing -K to any other 16 bytes changes the output completely (the key decides)

The three output shapes are the boundary: encoding changed the representation, hashing discarded information, encryption locked it.


Takeaways

  • Encoding solves format compatibility — reversible by anyone, nothing to do with secrecy;
  • Hashing solves consistency verification — irreversible; safety depends on salting and slowness;
  • Encryption solves secrecy — reversible with a key; safety depends on key management and mode choice.

All three have local tools on this site: AES encrypt/decrypt, text hash, Base64 converter — verifying these conclusions uploads nothing.

Advertisement

Frequently Asked Questions

Can Base64 be used to "encrypt" passwords?

**No.** Base64 is an encoding: no key, a public alphabet, and anyone can reverse it in one step. Storing `admin123` as `YWRtaW4xMjM=` in a database is identical to storing plaintext — "looks like gibberish" does not mean "confidential". Passwords belong in bcrypt/argon2 slow hashes.

If hashing is irreversible, why can it still be cracked?

Irreversible means you cannot derive the input from the hash — but attackers don't need to: they enumerate candidate passwords, hash each and compare. Bare MD5/SHA-256 password storage is unsafe because modern GPUs compute billions of hashes per second. The defence is **salt + slow hash** (bcrypt/argon2) to make each guess expensive.

The same file always produces the same hash — isn't that insecure?

Quite the opposite — **determinism** is exactly what makes hashes usable for integrity: identical input must give identical output, and a single changed byte produces a completely different hash (avalanche). "Security" depends on the use: SHA-256 is fine for verifying downloads; password storage needs salting and a slow hash, otherwise determinism just helps rainbow tables.

Is the Base64 in a JWT encrypted content?

No. A JWT's three segments are **Base64URL-encoded Header, Payload and Signature** — anyone can decode and read the first two (this site has a [JWT parser](/jwt-parser.html)). What prevents tampering is the third segment, the **signature**: modifying the Payload without recomputing the signature fails verification. So never put sensitive plaintext in a JWT payload.

← Back to Blog