Testing Files Generator
English

How to make a corrupt file for testing

A validator that has only ever been shown healthy files has not really been tested. Here is how to get a file that is broken on purpose, comes out at exactly the size you ask for, and carries a manifest saying what your system should do with it.

The short answer

tfg generate --format png --size 2mb --damage zero-head --out ./out writes a PNG of exactly 2097152 bytes whose first bytes are zeros, and the manifest beside it records that your system should reject it.

The usual way

Why a file corrupted by hand makes a poor test

The usual ways are a hex editor, a script that flips a few random bytes, or cutting a file short with head or truncate. They work once, and then they cost you:

What you get

A damaged file is still the size you asked for

The file is generated normally and broken afterwards, on its way to the disk. It keeps the size you asked for, and the same command writes the same bytes again.

tfg generate --format png --size 2mb --damage zero-head --out ./out
tfg generate --format pdf --size 1mb --count 5 --damage zero-head:bytes=16 --out ./broken

Settings go after a colon. The flag can be repeated, and the damages are applied in the order you write them. It works with every one of the 26 formats.

What it can do

Which damages are there?

This is the list the program prints, read from it when this page is built. tfg damage prints the same, and tfg damage <id> says what one of them takes.

Damage What it does to the bytes Smallest file Settings
zero-head Overwrites the first bytes of the file with zeros, leaving its length alone. Most readers look there first, so this is the damage almost anything notices. 8 bytes

zero-head writes zeros over the start of the file. Most readers look there first, at the signature and the header that say what the file is, so almost any reader notices. Plain text and logs have no signature and are turned away as well, because a run of zero bytes is not text. Below four bytes some formats come out with damage that no reader complains about, which is why the setting starts at four.

What the manifest says

A manifest that says what should happen

Every damaged file gets an entry saying your system should reject it, with the damage recorded beside it:

"expected": {
  "outcome": "reject",
  "reason": "content_malformed",
  "confidence": "certain"
},
"damage": [
  {
    "type": "zero-head",
    "settings": {
      "bytes": "8"
    }
  }
]

Two requests are refused before anything is written, because each would leave a file on disk that the manifest describes wrongly:

In a recipe

Healthy and broken files in one run

Put both in one recipe, and the manifest carries the expectation of every file, so the test does not need a list of which is which:

version: 1
targets:
  - id: healthy
    format: pdf
    size: 1mb
    expected: accept
  - id: broken
    format: pdf
    size: 1mb
    damage:
      - zero-head

In a test

Turning it into a test

The test reads the manifest and checks that what happened is what was declared. It needs no list of file names:

import json, os

directory = "healthy-and-broken"
manifest = json.load(open(os.path.join(directory, "manifest.json")))

for entry in manifest["files"]:
    response = upload(os.path.join(directory, entry["path"]))
    if entry["expected"]["outcome"] == "reject":
        assert not response.ok
    else:
        assert response.ok

A good refusal is a clean one. A message that says what was wrong is the answer you want. A server error, a hang or a half stored file is the defect this test exists to find.

Next

Where to go from here