Sep 20, 2026 · by Lukas · View source

slop-grader

Jev-AI CLI tool that evaluates text against custom rulesets

slop-grader

Editorial analysis

The real content problem in 2025 isn’t volume — it’s slop

If you run social accounts for a living, you already know the dirty secret of the AI era: publishing has never been easier, and standing out has never been harder. Every creator I know is now shipping three to five times the output they managed two years ago, because the tools got cheap and the platforms got hungry. But the feed is drowning in the same synthetic cadence — the “in today’s fast-paced world” opener, the tricolon, the em-dash pivot, the emoji-sprinkled listicle that says nothing. Audiences are starting to smell it, and the algorithms are starting to penalize it, because dwell time and saves are the metrics that actually matter now, not raw impressions. So when a tool shows up that tries to grade text for AI filler before it goes out the door, I pay attention. That’s what slop-grader is — an open-source linter for prose, built by Lukas on top of a model called Jev from TypeSafe. It’s niche, it’s developer-flavored, and it is not a social media scheduling tool. But the underlying idea — codifying your editorial standards as machine-checkable rules — is one of the more useful mental models I’ve seen a creator tool ship this year, and I want to unpack why.

What slop-grader actually does, stripped of launch-page language

The maker’s own description is refreshingly plain: it’s an open-source tool that checks any text document against a set of rules. Out of the box you get English grammar, German grammar, and AI filler detection. Beyond that, you write custom rules in plain language — SEO checks, legal clauses, tone-of-address consistency (his example is keeping German “Du” versus “Sie” from drifting mid-document). The rules come in two flavors, and this distinction matters more than it sounds: some are evaluated line by line (“Does this line make a promise that requires a legal disclaimer?”) and some are evaluated across the entire document (“Does the opening earn the reader’s next 30 seconds?”). The tool outputs a list of flagged lines plus instructions you can paste directly into an AI agent to fix the document — that last part was confirmed by the maker in the comments when a user asked whether the output was meant for manual review or agent workflows.

The engine underneath is the interesting bit. Jev is described as a “System One” model — explicitly not an LLM — specialized in answering structured yes/no questions. That’s why checking a document reportedly takes seconds and costs less than a cent. To run it locally you need Node.js and an account with either TypeSafe or OpenRouter, and the maker flags clearly that text is evaluated on an external AI server. No pricing tiers, no SaaS dashboard, no free trial funnel — not disclosed because there isn’t one. It’s a CLI-adjacent utility, and you should treat it as such.

Why a “linter for prose” is the right metaphor

If you’ve ever worked in a codebase, you know what ESLint does: it doesn’t write your code, it just tells you where you violated a rule you already agreed to follow. Slop-grader is that, but for sentences. And that framing solves a problem most social teams have never named out loud. Your brand voice doc says “we never use exclamation points” or “we always lead with the customer’s problem.” Nobody reads it. Nobody enforces it. Six months later your LinkedIn posts sound like a different company than your TikTok captions. A ruleset that runs on every draft is the enforcement layer your style guide never had.

How this differs from the AI-detection tools you’ve probably tried

Let me be blunt about the category, because there’s a lot of garbage in it. The incumbent “AI detector” space — things like GPTZero, Originality.ai, Copyleaks — is built around a fundamentally different question: “Was this written by a machine?” That’s a forensic question, and it’s a losing one. False positives are rampant, the classifiers are perpetually behind the models, and for a social media manager the answer doesn’t even matter. What matters is: does this copy read like something a human would actually want to read?

Slop-grader sidesteps that trap entirely. It doesn’t care who or what wrote the draft. It cares whether the draft violates rules you defined. That’s a much more defensible position, and it’s closer in spirit to editorial tools like Grammarly or Hemingway than to detector SaaS — except those tools ship a fixed ruleset you can’t meaningfully extend, while slop-grader’s entire value proposition is that you build the ruleset yourself.

The other comparison worth drawing is against the AI writing assistants themselves — Jasper, Copy.ai, Writesonic. Those tools generate. Slop-grader critiques. In my experience, the generation side of the stack is now commoditized and the critique side is where the leverage is, because the bottleneck for a working creator isn’t producing a first draft anymore — it’s catching the tells before you hit publish.

Where the math breaks

Here’s my honest read on the economics. The maker says a document check costs less than a cent and takes seconds. If that holds up across a real workload — say 40 caption drafts, 12 long-form LinkedIn posts, and a handful of YouTube scripts a week — you’re looking at pennies per month in model costs, plus whatever TypeSafe or OpenRouter charges on top. Compare that to the seat-based pricing of a Buffer or Hootsuite add-on and the per-token costs of running every draft through a frontier LLM with a big style prompt. The math is genuinely favorable, and it’s favorable because it’s not an LLM. That’s the whole trick.

The catch: you have to write the rules. And writing good rules is a skill. Which brings me to the section most launch coverage skips.

What creators and social teams can actually borrow from this

Even if you never install Node.js in your life, there are three operational ideas here worth stealing this week.

1. Turn your brand voice into yes/no questions, not adjectives

Most style guides are written as adjectives: “bold, warm, human, irreverent.” Adjectives are unenforceable. Questions are enforceable. Rewrite your guide as a list of binary checks — “Does the first line name a specific pain the audience has felt this month?” / “Does any line claim a result without a caveat?” / “Does this post use a word our competitors also use in their last ten posts?” That last one is a killer rule for anyone doing competitive positioning on LinkedIn or X, where the feed is a wall of identical thought-leadership cadence.

2. Split your rules into line-level and document-level

This is the sharpest design choice in the whole product, and it maps directly to how social copy fails. Line-level failures are tactical — a banned word, a missing disclosure, a CTA that doesn’t match the platform. Document-level failures are structural — a hook that doesn’t earn the scroll, a caption that buries the payoff, a thread that peaks in tweet three. Most social teams only review at the document level, which is why they miss the small stuff, or only at the line level, which is why their posts are technically clean and strategically dead. You need both passes, and they should be separate passes.

3. Route the output to an agent, not a human, for the first fix

The maker confirmed the flagged lines and instructions are designed to be pasted into an AI agent. That’s the workflow most creators are already half-running — draft in Claude or ChatGPT, edit by hand, schedule in Later or Metricool. Slop-grader slots in as the QA step between draft and edit, which is exactly where human attention is most expensive and most wasted. Let the machine catch the “in today’s fast-paced world” opener. Save your brain for the hook rewrite.

Why TikTok and YouTube creators should care more than LinkedIn ones

I’d bet the ROI here is inverted from what you’d expect. LinkedIn is where AI slop is most tolerated — the audience is scrolling at work, half-reading, and the algorithm rewards consistency over craft. TikTok and YouTube are the opposite: watch time and retention are brutal, unforgiving metrics, and a script that opens with three seconds of throat-clearing gets scrolled past before the algorithm even registers the impression. If you’re scripting short-form video in CapCut or writing long-form YouTube intros, a document-level rule like “Does the opening earn the reader’s next 30 seconds?” is worth more than any grammar check. The German “Du” versus “Sie” example the maker cites is a localization tell, but the same principle applies to register drift across a 20-minute video — the tone shifts, nobody notices, and the audience quietly leaves.

Where I think it falls short — and who should skip it

Let me be transparent about the limitations, because the launch page doesn’t volunteer them and the comment thread surfaces a few.

It doesn’t explain its reasoning. A user named Gal Dayan asked directly whether the tool explains why a rule fired, or just returns the line number and rule text. The maker’s answer: it does not give you the reason, but you can write separate rules for each check and the agent is good at inferring the problem. My take: that’s a real friction point. A linter that says “line 14 violates rule 7” without a rationale is a linter you’ll learn to ignore, especially on a team where junior writers need to understand the fix, not just apply it. The workaround — atomizing rules — is legitimate but shifts cognitive load onto whoever writes the ruleset.

The strictness question is unanswered. Another commenter, Amelia, suggested a strictness setting so users could toggle between quick cleanup and deep review. The maker didn’t respond to that one in the thread. As of the launch, no such setting is described. So you get one intensity level, and you tune it by adding or removing rules. Fine for power users, annoying for everyone else.

It’s not a social media tool. No scheduling, no platform-native previews, no character-count awareness for Threads or Pinterest descriptions, no image or video analysis — it’s text-only. If your bottleneck is visual content or posting cadence, this does nothing for you. If your bottleneck is that your captions and scripts all sound the same, it might be the highest-leverage sub-cent you spend this quarter.

The setup tax is real. Node.js, a TypeSafe or OpenRouter account, and external AI server evaluation of your text. For a solo creator comfortable in a terminal, that’s an afternoon. For a social team inside a larger org, the “text is evaluated on an external AI server” line alone will trigger a legal review, and rightly so — you don’t want to paste client NDAs or unreleased campaign copy into a tool without a clear data-handling story. The maker is upfront about this, which I respect, but it’s a genuine adoption blocker for agency and enterprise workflows.

It’s open source, which cuts both ways. You can inspect it, fork it, and extend it — the custom-rule creation skill lives in the project’s GitHub repo. But open source also means no SLA, no support desk, and a maintenance burden that depends entirely on one person’s continued interest. That’s fine for a personal workflow. It’s not fine as the enforced QA layer for a 12-person content team without someone owning the fork.

Who this is NOT for

If you’re a visual-first creator whose captions are three words and an emoji, skip it. If you’re a social media manager who needs a unified inbox, approval workflows, and a client-facing analytics dashboard, this is not in your category — look at Sprout Social or Agorapulse instead. And if you don’t have a written style guide yet, the tool has nothing to enforce. Write the guide first. The ruleset is downstream of the strategy, always.

What I’d watch / test next

Three concrete moves for this week, in order of effort-to-payoff.

First, spend 30 minutes turning your existing brand voice doc into 10 to 15 yes/no questions. Don’t install anything. Just write the questions and run your last five published posts against them by hand. I’d bet you catch at least two that violate rules you didn’t know you had. That exercise alone is worth the price of admission.

Second, if you’re technically inclined, clone the repo and run slop-grader against your next batch of drafts before they hit your scheduler. Track two things for a month: how many flags you actually act on, and whether engagement rate on the flagged-and-fixed posts diverges from the unflagged ones. That’s the only honest way to know if the ruleset is earning its keep — and it’s the kind of experiment almost nobody in the creator tooling space runs, because everyone’s too busy chasing the next shiny dashboard.

Third, watch the rule-sharing ecosystem. If slop-grader’s custom rulesets become shareable — a “newsletter voice pack,” a “B2B LinkedIn pack,” a “short-form script pack” — that’s when this stops being a developer curiosity and starts being infrastructure. The maker hasn’t announced anything like that, so I’m not predicting it. But it’s the obvious next move, and it’s the one I’d be watching for. Until then, the leverage is in writing your own rules, and that work is yours no matter which tool you use.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with FLOWNIB. No editing skills required.

Start Creating for Free