Tribar Studio
← Back to blog
AIEngineering Culture

Anti-slop engineering: the checks we run before model output ships

Engineering Team·

Every team shipping AI features discovers the same uncomfortable truth: a model that produces something plausible is not the same as a system that produces something right. Plausibility is exactly what these models are good at. Correctness is what production needs. The gap between them is where slop lives.

Slop is not a model problem. It is a review problem. The model will always generate; the question is what your pipeline lets through.

Here is the pipeline we run before anything model-produced reaches a user.

1. The output must fit a schema, not a vibe

If a feature's output will be consumed by code, it does not get to be prose. Structured outputs — JSON schemas for API responses, typed tool calls for actions, enums for classifications — move failure from runtime interpretation to validation time. A model that returns "severity": "fairly urgent" where the schema says enum ["low", "medium", "high"] fails at the boundary, visibly, instead of three layers deeper as a silent misclassification.

We reject prose pipelines for machine consumers the same way we would reject an API that returns "some number, probably".

2. Deterministic checks before model-based checks

The cheapest eval is a function. Counts, ranges, required fields, cross-references, unit conversions, date sanity — anything checkable with code runs first, because it is fast, free, and never hallucinates.

generate → schema validation → invariant checks → model-based eval → human review

Only what survives the deterministic layer deserves the expensive evaluation. Most bad output dies at step two, and it should — spending model calls to grade garbage is its own kind of slop.

3. Evals pinned to decisions, not to vibes

An eval that says "the summary felt good" is theater. Evals we ship with have to encode a decision: did the classifier pick the right category, did the extraction miss a required clause, did the translation preserve the numbers. They run in CI like any other test, on a fixed corpus that grows with every incident.

The corpus discipline matters more than the eval framework. Every production miss becomes a permanent test case — that is how the system gets quieter over time instead of louder.

4. A human path that actually exists

Regulated and human-facing surfaces need human oversight, and oversight that only exists on a slide is non-existent. We ship the boring version: confidence thresholds that route low-confidence output to a human queue, UI affordances that make review the path of least resistance, and audit trails that record who approved what.

The test question is simple: when the model is wrong at 2 a.m., is there a mechanism that catches it, or a hope?

5. The fallback is part of the feature

Every AI feature we build works when the model does not: the search degrades to keyword matching, the triage degrades to rules, the draft degrades to a template. This is not pessimism — it is the difference between a feature that has a bad day and a feature that has an outage.

Fallback design also keeps teams honest. If the fallback path produces almost the same value, the model might not be earning its place. If it produces nothing, the feature is load-bearing on a non-deterministic dependency — better to know now.

What we actually optimize for

The Human Hand test is our internal shorthand: would a model averaging the internet produce this output? If yes, the work is not done. That test applies to the checks above too — a pipeline that "checks quality" with another model, another prompt, and another cloud of best intentions is still slop, just one layer up.

Deterministic gates, decision-shaped evals, real human paths, honest fallbacks. None of it is glamorous. All of it is what "AI features you can trust" reduces to — the model supplies capability; the pipeline supplies trust.