fae
All work

Fae · specimen №04

FluffFilter

Whether a piece is good enough to publish, and exactly what to fix if it isn't.

A first-pass editor for content teams: upload a document, or fifty, and get an accept, revise, or reject verdict, scores across five quality dimensions, and three to five surgical fixes with exact locations and replacement text.

Role
Product design, engineering, evaluation design
Platform
Web
Years
2025–present
Built with
Ruby on Rails · Hotwire · Tailwind CSS · Language-model evaluation pipelines
Status
Live
The FluffFilter homepage: AI made content cheap. Quality didn't get the memo.
From the app.

The problem

Content got cheap to produce. Quality didn't get the memo. A marketing team's output grows tenfold while its review bandwidth stays flat, so editors spend their days doing copy triage on a conveyor belt of plausible-sounding drafts instead of the strategy work they were hired for.

The tools that exist answer the wrong question. A detector tells you whether a machine wrote the words, which matters far less than whether the words are any good. And feedback like "make it more engaging" leaves the writer exactly where they started, only more discouraged.

The approach

FluffFilter reads a document the way a ruthless editor would and reports the way a kind one does: a verdict (accept, revise, or reject), a score out of five, and a short list of fixes, each with the exact location, what's wrong there, and replacement text ready to paste.

A sales email and a PRD fail differently, so there is no one-size-fits-all rubric. More than twenty content types each carry their own evaluation criteria, and the type is detected automatically. The framing is deliberately modest: a junior editor that never gets tired, doing the first pass so the humans keep the brand voice, the creative direction, and the final word.

Decisions that mattered

4 entries

Not another detector

Whether a machine wrote a draft is the least useful question you can ask about it. FluffFilter answers whether it should ship, and what to change if not. Detection assigns blame; evaluation improves the work.

Every fix has an address

Findings name a location, state the problem, and offer replacement text. "Add social proof" just relocates the triage work onto the writer; a fix you can copy and paste actually removes it.

One rubric per content type

A landing page fails by burying its ask; a PRD fails by leaving a requirement ambiguous. Grading both against the same checklist flatters one and flunks the other, so every content type gets its own evaluator.

Documents are not training data

A content team's unpublished drafts are its business. Uploads are analyzed and returned, never used to train models. For a tool that reads your work before the public does, saying so plainly, in the FAQ rather than a policy PDF, is table stakes.

What it does

In a review

  • An accept, revise, or reject verdict with a score out of five
  • Scores across five quality dimensions
  • Three to five surgical fixes: location, problem, replacement text
  • Content type auto-detected against 20+ specialized evaluators

For the team

  • Batch uploads: fifty posts reviewed in minutes
  • Full analysis history
  • Plans from solo starter to agency scale, each with a 7-day trial
  • Higher plans choose exactly which evaluators run

Marginalia

FluffFilter is the commercial bet in the catalog. The flood of machine-written content isn't receding, which makes judgment the scarce resource. This is that judgment, sold by the analysis.

More figures

From the app

A FluffFilter report: verdicts on value proposition, social proof, and call to action, above a prioritized list of fixes.
Plate I: The report. Critical gaps first, then fixes ready to paste.

The rest of the catalog

All work