Text encoding
Does my reader know which encoding a file is in, or is it guessing?
The text-encoding preset builds a whole set of real test files for this question in one
command, and a manifest.json beside them saying how your system should react to each
file. Everything below is read from the program, at the defaults of this version.
What does it usually find?
- a reader that assumes UTF-8 and shows a UTF-16 file as one character in three, or as rows of boxes
- a byte order mark read as content, so the first field of an import starts with three stray characters
- an importer that guesses the encoding from the opening bytes and guesses differently for a longer file
- a CRLF file split into rows with an empty row after each one, or a carriage return kept inside the last field
What is in the set?
At its defaults, as tfg preset show text-encoding reports it:
| Files | 20 |
|---|---|
| Targets in its recipe | 20 |
| Total size | 81 920 B |
| Formats | csv, log, md, txt, xml |
And what the manifest of that set expects from your system:
| Expected | Meaning | Files |
|---|---|---|
accept | Your system should take the file. | 10 |
unspecified | It depends on the rules of your system. You decide, then check that what happens is what you meant. | 10 |
What can you change?
| Setting | Takes | Default | What it does |
|---|---|---|---|
--sample |
a size such as 2mb | 4kb |
How big each file of the set is. UTF-16 stores two bytes for every character, so an odd number is refused. |
How do you run it?
See what the set would cost, build it, or take its recipe to edit:
tfg preset show text-encoding
tfg generate --preset text-encoding --out ./text-encoding
tfg preset eject text-encoding > text-encoding.yaml
Or build on it in a recipe of your own, next to your tests:
version: 1
extends: preset:text-encoding