The most important shift in the creator economy isn’t that AI can write captions, cut clips, and schedule posts. It’s that AI can do all of that for everyone — which means production is no longer the scarce skill. Review is. The ability to look at an artifact you didn’t create, find what’s broken, and steer it to done under a strict budget of revisions and tokens is the core competence of any social media operation that uses AI seriously. So the most useful thing I read on Product Hunt this week wasn’t a social tool at all. It was Merge, a hiring platform for engineers that scores candidates on how well they review code, not how well they write it. Read sideways, it’s a blueprint for how content teams should evaluate themselves in 2025.
The old test measured the wrong thing
When I hire freelancers to write captions or edit short-form video, my first ask is a portfolio. For the past eighteen months, I’ve had a suspicion every social media manager shares: half the work I’m reviewing was assembled with heavy AI assistance. It’s becoming impossible to tell whether a strong portfolio reflects the operator behind it. For a hiring manager, that’s a measurement crisis. For a creator, it’s an existential one — if your portfolio can be faked, so can everyone else’s, and the market will quietly reset the value of portfolio work.
Harshith Latchupatula, the maker of Merge, describes the same breakdown on the engineering side: five founders, collectively more than 250 interviews across startups, FAANG, and quantitative trading shops, and a blunt conclusion — LeetCode-style interviews test skills that weren’t used on the job. Even the AI-assisted interviews, the maker says, still grade code output as the primary metric. Meanwhile, at their own companies, pull-request counts have nearly tripled, and most of the founders say they haven’t manually edited a line of code in a year. The founders’ argument is that code production will not remain the hardest engineering skill; the real difficulty is reviewing code that another engineer’s AI generated, with very little context of your own. The work moved from writing code to reviewing a stream of code written by humans and AI agents. The hiring process never moved with it.
Swap “engineers” for “creators” and the parallel is exact. Your feed is a stream of posts written by humans, edited by AI, or both. The value is not in the first draft; it’s in the second, third, and fourth passes. It’s in catching a brand-safety problem in a sponsored caption before it ships to 50,000 followers, and telling the AI exactly what to fix without burning five regenerations on a single post. Those are review skills, not production skills. And almost nothing on the market trains or measures them.
The scoring rubric is a content workflow in disguise
Merge’s flow: a candidate is shown a small codebase and a pull request, then reviews it the way a real engineer would comment. An AI agent responds to each comment with a code change or a reply, simulating a human teammate. The loop can repeat until five revisions are used or time runs out. The platform scores three dimensions: coverage, communication, and efficiency. Those three words deserve a permanent home in your content operations.
Coverage, in Merge’s terms, is how many bugs or vulnerabilities the candidate identified and addressed. In content terms, it’s which known failure modes you check before hitting publish. Many teams ship 30 posts across five platforms a week and could not write down, in a sentence, what “good” means for each format. A coverage checklist for a sponsored reel might read: brand-safety scan, CTA alignment, hook strength in the first 1.5 seconds, every factual claim supportable, caption matched to format conventions. The metric is simple: how many of your known failure modes did you catch before publishing?
Communication is whether the feedback was efficient and constructive. In my own review workflows, vague editorial feedback is the largest time-waster in an AI-assisted pipeline. “Make it punchier” burns a revision, generates a new draft, and burns another token. “Open with the statistic, move the context to the second sentence” is efficient and constructive. That is a human skill, and it is the skill a senior creative operator now uses all day: directing AI without giving up taste.
Efficiency is the part that should make you uncomfortable. The Merge team claims the platform is the first to show how a candidate uses tokens, LLM costs, and pull-request revisions. That’s engineering-specific, but apply it to content: every regeneration is a cost line. If a single post takes four AI drafts and three human edits, you’ve spent four times the token budget and three times the editorial attention on an asset that might perform the same as one shipped in two revisions. In my own workflow, I’ve standardized a budget: two generated drafts per asset, one human revision, done. If it’s not publishable after that, it goes in the bin. That discipline is the direct descendant of Merge’s five-revision cap — and it’s the only thing preventing AI-era content from drowning in infinite polish. The reason this matters doubly for solo founders is that you often have no review layer at all. If an indie founder generates a month of content in a day, the review pass is either skipped or done by the same person who is too close to the material. Merge’s design forces a second party into the loop — an agent that pushes back, revises, and makes the reviewer account for their feedback. That is a workflow pattern worth borrowing even if the product never ships for your use case.
Why short-form creators should care more than long-form writers
If you publish long-form essays, the review shift is real but gradual: the written word rewards a point of view, and AI still struggles to produce strong opinions, so the human layer is easy to detect. Short-form video is a different universe. TikTok’s algorithm distributes on retention, the first 1.5 seconds decide whether a post is seen, and a slightly mistimed caption edit can tank watch time. AI can generate a script and a rough cut in minutes, but the cost of shipping something sub-optimal is immediate — you miss the trend cycle and the distribution window. In that world, the person who can review a draft edit, spot the retention killer, and direct a targeted fix under a strict revision budget is worth ten people who can generate drafts. Merge’s rubric applies to short-form video production more directly than to any other content vertical I can think of.
The incumbents in this category are generation-focused platforms like LeetCode: they ask “can you produce a solution” and grade the artifact. Merge asks a harder question — “can you look at someone else’s artifact, spot the breakage, and drive it to done without exhausting the budget?” The content-industry equivalent is the gap between scheduling and reviewing. Buffer and Hootsuite measure what you published and when. None of them measure judgment: the ability to look at a first draft, find brand risk, spot the hook that will flop, and fix it conclusively. The promoted slot at the top of the launch page tells the same story: Framer AI Agents can now design and publish a professional site with minimal human input. Production cost for a website, a post, or a video is heading toward zero. The scarce unit is the person who reviews, corrects, and ships — in engineering and in content alike.
What you can steal from Merge this week
Five transfers.
First, write your coverage list. Not a generic brand-voice checklist — a failure-mode list specific to each content type. For a thought-leadership post: is the opinion defensible? Does the first line earn the second? Could the claim be clipped out of context and embarrass the brand? For a community reply: does it escalate or de-escalate? For a product teaser: does the hook match what the product actually does? You cannot measure coverage if you haven’t defined the failure modes.
Second, cap the AI revision loop. Merge caps candidates at five revisions; I cap drafts at two. The constraint does not reduce quality. It forces a decision instead of another pass. In my experience, a team that ships “good enough, on-brand, now” beats a team that ships “perfect, generic, Thursday.”
Third, count revisions per asset. Track drafts-to-publish and human-edit passes for each format. You’ll find, as I did, that certain formats are revision sinks. The fix is usually not a better model; it’s a better brief or a clearer reviewer. The number you’re looking for is not “how many posts you published” but “how many iterations the average post consumes.” That is your true cost of content.
Fourth, grade your own feedback for constructiveness. “Too wordy” is an aesthetic opinion. “Cut the second paragraph; it repeats the hook” is feedback that reduces the next revision’s cost.
Fifth, put a machine reviewer inside your pipeline — not at the end. The maker’s finding is instructive: an AI reviewer catches a lot, but its recommended fixes create new problems, and low revision efficiency is a signal of weak ability because it increases merge latency in real life. In content terms: an AI proofreader will flag passive voice and never notice that you’ve drifted off brand voice. Put the AI pass after the first draft, before the human full pass, and watch whether it lowers the number of revisions the human must do. If it just adds a round, your tooling is wrong. This is also why I tell social teams not to replace junior editors with AI, but to give each junior editor an AI critic that argues back. The friction is the feature — a tool that always agrees with you teaches you nothing.
Where the math breaks
Now the part the launch page won’t tell you.
Merge is a pitch, not a proof. Pricing is not disclosed. Candidate volume is not disclosed. The team asks visitors to book a demo, which suggests they are early in customer discovery. The most honest moment on the page is the comment section. Artem Fedorovich asks how Merge scores the final diff versus the reasoning behind it, and more pointedly, how it keeps the exercise from rewarding whoever has the better model open in a second tab. The maker’s answer: both are scored, with a heavy bias toward reasoning, because AI-only reviewers fix one issue and create another. That’s plausible. It is still an arms race. Omri Ben-Shoham spells out the end state: candidates run the take-home through their own AI code-review tool before submitting, so eventually it’s AI grading AI.
Here’s my skepticism. If the candidate’s environment is unmonitored, token efficiency and revision counts can be gamed by anyone who knows the scorecard. Write minimal cheap comments, use a small model to look efficient, and let the real analysis happen off-screen. The metrics inherit the classic dashboard failure: once you turn efficiency into an exam question, people optimize the metric instead of the outcome.
The arms race nobody can close
The deeper problem is that token use and revision count are cost metrics, not quality metrics. A candidate who catches five issues in two revisions is efficient. A candidate who catches zero in one revision looks identically efficient if you only read the revision column. Coverage is supposed to catch that, but coverage depends on the test being seeded with the right number and severity of bugs — and candidates will learn the average bug density the way they learned LeetCode patterns. I’d bet the next arms race is candidates sharing the specific Merge fixtures. The content-world version is already here: if you announce a “revision cost” KPI for your team, the unambitious will optimize the score by doing less, and the ambitious will quietly do the work elsewhere and log a clean score. Reward the outcome — the quality of the final review, the post-market performance, the brand-safety record — not the dashboard.
Who Merge is not for
Merge looks well aimed at senior engineering hires who will live in a codebase full of AI-generated code. I don’t see it replacing take-home assignments for junior production roles: a candidate with low context but a sharp eye reveals little across five revisions. The same boundary applies to content. If you hire a junior social coordinator whose job is mostly scheduling, clipping, and posting, you need a speed-and-accuracy test, not a judgment assessment. Judgment assessments belong to roles where taste compounds: lead editors, brand-voice owners, and strategists who tell the AI what to make and the humans how to fix it.
Be honest about the load-bearing assumptions too. The 250-interview figure and the tripled PR counts come from the founders’ own experience, not from an independent study. The scoring rubric has not been validated against on-the-job performance, and no calibration data is published. A confident hypothesis is still more than most tools ship with — but it is not yet a benchmark.
What I’d watch / test next
Here’s what I’d watch: whether Merge publishes a calibration post — a transparent rubric and a dataset showing how review scores correlate with actual engineering performance. If it comes, read it the way you read an algorithm update: does the metric predict the outcome, or just itself?
Here’s what you can test this week, without Merge. Run a coverage audit on your ten most recent posts: list the failure modes you checked before publishing, then the ones you missed. That baseline tells you where your checklist needs teeth. Then cap the AI loop at two regenerations per asset for seven days, and count how many times you hit the cap. My bet: the constraint surfaces more judgment in your review pass, because you finally have to decide instead of requesting one more draft.
If you want to track Merge’s scoring model as it matures, follow the team on X and LinkedIn, and read the maker’s replies on the launch page. The whack-a-mole description alone is worth the visit — even if you never plan to hire an engineer.




