PolyPress

Principles · why both tools are built the way they are

Measure it,
or don't claim it.

Two tools that do entirely different things were built under the same set of rules. The rules are the transferable part.

None of this is a philosophy of software in general. It is a narrow response to one specific failure mode — the tool that returns a plausible answer without anyone able to tell that it is wrong — and every rule below exists because that failure actually happened during the work.

01

Eight rules

  1. Never worse — measured, not assumed

    In Polypress, a modelled encoding must beat the plain fallback, and a parent column must beat no parent. Every guard works by encoding both ways and keeping the smaller. The guarantee is not reasoned about; it is executed. This is also what makes the compressor slow, and that trade was made deliberately.

  2. Verify before writing

    Both Polypress command-line tools decode the compressed blob and compare every cell against the input before any file is created. Nothing lands on disk that hasn't already been proved to round-trip.

  3. Entropy is a good nominator and a bad decider

    Use a cheap statistic to shortlist a choice, then actually compress and compare. Fixing one place where this rule was broken took 19.7% off the codec's best dataset. The same rule, arrived at independently, is written into the expired AT&T patent that covers the core idea.

  4. A component probe may nominate; only a whole-file encode decides

    Four separate transforms measured +12.7%, −24.9%, 0.647× and −1.02% in isolation — and three of the four made the whole codec worse. A column is not worth its own bytes; it is worth its own bytes plus its value as a sort parent for every other column. Local improvement is not evidence of global improvement.

  5. Every all-or-nothing test is suspect

    Four blank cells in 72,048 cost 41.2% of a Treasury yield curve, because a single failing cell disqualified a whole column and with it the whole 2D predictor. One -0.0 in 26,304 cells split a 21-column matrix into 18 and 3.

  6. Treat your input as hostile

    The Polypress decoder reads files other people made. Corrupt input must be refused, never crash, never allocate unbounded. There is a dedicated test suite of deliberately lying headers.

  7. Nothing is hidden, and nothing is a private helper

    Handrail's rule: no generated line may call a function that isn't documented in R or in a named package. Defaults that a beginner would be surprised by — na.rm, fixed, row.names — are written out rather than relied on, so the code on screen matches the documentation they will go and read.

  8. A choice a researcher must defend cannot be made silently

    If a method would appear in a paper's methods section, the tool does not get to pick it. This is why Handrail refuses gtsummary and tableone for Table 1 — both are better-looking, and both choose a statistical test per variable by their own rules. Plain R is the floor because it is the only option that chooses nothing.

02

Negative results get written down

The Polypress README carries a table of everything that was tried and failed, and it is treated as the most valuable part of the document. Handrail's brief lists the five standard Swift mistakes that made a 278 MB file take two minutes to open instead of four hundredths of a second.

The reason is not modesty. A record of what failed is the only thing that stops the same idea being re-attempted, and it is the only part of a result that tells a reader how hard the author actually looked. One of the five Swift mistakes — the missing autoreleasepool, worth 304 MB against 21 MB — took the same wall-clock time either way. No amount of reasoning would have surfaced it. Only the measurement did.

The same rule applies to claims that turned out to be wrong. Polypress originally claimed its central idea was novel; a prior-art search found a 2009 patent disclosing the whole of it. The claim was withdrawn, and the withdrawal is on its page rather than quietly deleted.

03

Human in the loop, precisely

“Human in the loop” usually means a person approving an output they cannot evaluate. That is a signature, not a check. For it to mean anything, three things have to be true at the moment the person is asked to look.

The thing shown must be the thing that runs. Handrail shows you dplyr and then writes that same dplyr into your file; there is no separate internal representation, no translation step, nothing that could drift. You run it yourself, in RStudio, with Cmd-Enter.

It must be checkable with ordinary skill. Every function on a generated line is documented in R or a named package, so verifying it is a search away rather than an appeal to the tool's authority.

The tool must state where it is uncertain. The survey-weighting problem on Handrail's page is unresolved and named as unresolved, because the alternative — generating naive weighted tests — produces numbers that look completely fine.

The point is not that a person is present.
It is that a person could tell if it were wrong.
04

What this is not

  • Not an argument against automation

    Both tools are automation. Polypress makes thousands of modelling decisions a person could never make by hand; Handrail writes code so its user doesn't have to memorise syntax. The objection is not to a machine making the choice — it is to the choice being unavailable for inspection afterwards.

  • Not “no code”

    Handrail is explicitly not a no-code tool and not “R without code.” It hands over R's own vocabulary, one line at a time, and the generated script runs with or without the app installed.

  • Not a claim of novelty

    Polypress's core idea is prior art, independently reached. The contribution is a careful, measured, verified implementation with a table-aware front end — tested harder than the prior art was, not thought of first.

  • Not finished, and not independently confirmed

    Nobody outside these projects has run either tool. Every benchmark dataset is a government open-data table. Handrail has no icon, no installer, and its most-wanted features are unbuilt. Those facts belong next to the results, not behind them.