Aug 31, 2026 · by Peder Aaby · View source

Blume.codes

Turns coding agent sessions into better rules and skills

Blume.codes

Editorial analysis

Why a Tool That Watches Your AI Assistant Matters More Than Another Scheduling App

Every social media operator I know is running the same experiment right now. We’ve handed our content calendars, our brand voice guidelines, and our repurposing workflows to AI assistants, hoping they’ll turn a 45-minute task into a five-minute one. And for the first few weeks, it works beautifully. The AI drafts captions that sound like you. It formats the YouTube description correctly. It remembers that your audience hates emoji spam on LinkedIn but tolerates them on Instagram.

Then, quietly, things start to rot. The AI forgets the tone shift you corrected it on three weeks ago. It starts defaulting to the same hook structure you explicitly told it to stop using. You find yourself typing the same correction into the session prompt for the fifth time, wondering why the system doesn’t just learn from the fact that you’ve now said “no, shorter, punchier, and stop with the hashtag blocks” in four different ways across two weeks.

This is the exact problem that caught my eye when I saw Blume.codes launch on Product Hunt. It is not another scheduling dashboard. It is not another AI content generator. It is a desktop app that sits next to your coding agents—Claude Code, Codex, Cursor—and watches your sessions for patterns of correction, frustration, and repeated nudges, then clusters them and suggests concrete updates to your agent’s setup. The team behind it, Blume.codes, built it after their previous startup suffered from what they call “agent drift”—the gradual decay of an AI’s usefulness as your project grows and its context becomes stale.

Now, you might be thinking: I don’t write code, I write captions. Why should I care about a developer tool? Because the underlying mechanics are identical to what happens when you run a content operation with AI assistance. The same drift that hits a codebase hits a content system. The same frustration you feel when your AI writing assistant ignores your brand voice is the frustration a developer feels when their AI coding assistant ignores their architecture preferences. And the solution Blume is proposing—extracting human intent from sessions and using it to improve the system—is exactly what every social media manager needs but doesn’t have a tool for yet.

Let me explain why this matters, what it actually does, and where I think it falls short.

The Real Problem: AI Context Decay Is the Silent Killer of Content Operations

Here is what nobody tells you when you start using AI for content production: the tools are excellent at the first session and progressively worse at every session after that. When I set up a new AI assistant for a client’s social channels, I spend the first day writing detailed instructions. I paste in their brand guidelines, their tone of voice examples, their do-not-say list, their content pillars. I create a system prompt that would make a corporate lawyer proud. The first ten outputs are genuinely impressive.

Then the client sends feedback. “We don’t say ‘delve’ anymore, it sounds corporate.” “Our audience is Gen Z, stop using ‘utilize.’” “The CTA needs to be more urgent, we’re running a sale.” I correct the AI in the session. It produces a better caption. I move on. Three days later, I’m in a new session, and the AI has completely forgotten the “delve” correction. It’s back to writing like a mid-tier consulting firm. I correct it again. This cycle repeats until I finally spend an hour updating the system prompt, hoping I remember every single correction I’ve made across the last ten sessions.

This is agent drift, and it is not a coding problem. It is a context management problem that affects anyone who uses AI assistants for ongoing, iterative work. The maker of Blume, Peder Aaby, describes the failure mode precisely: “Duplicated functions, incoherent architecture and sneaky production bugs” in code. In content terms, that translates to duplicated brand voice instructions scattered across multiple documents, incoherent content pillars that contradict each other, and sneaky tone violations that slip through because the AI’s context window is a mess.

The current best practices for fighting this drift are manual and they rot. You try to maintain proper rules, skills, and documentation for your AI assistant. You ask the AI to self-verify its outputs against your guidelines. You think you’ve got it covered. But as your content library grows—as you produce more posts, more scripts, more newsletters—the setup rots. Manual maintenance becomes too time-consuming. Asking the agent to maintain its own context leads to what Aaby calls “context drift and bloat.” The AI starts paying attention to the wrong signals, or it accumulates so many instructions that it can’t prioritize them effectively.

What Blume does differently is extract the signals from your sessions that you’re probably not even aware you’re sending. Every time you correct the AI, every time you nudge it in a different direction, every time you type a frustrated ALL CAPS message because it just generated the wrong thing again—those are data points. Blume clusters them thematically. If a cluster reaches a pain threshold or a recurrence threshold, it suggests a concrete update to your setup. That could mean updating a stale rule, patching a hook or skill, or creating an entirely new one. You review the suggestion and apply it. The agent improves for the next run.

In my experience testing similar tools and running my own AI-assisted content workflows, this approach is fundamentally sound. The signals are there. The problem has always been that we don’t have a systematic way to capture them and turn them into durable changes. We rely on memory, which is unreliable, or on manual documentation, which is tedious. Blume is attempting to automate the capture and the clustering, which is the right instinct.

Why TikTok Creators Should Care More Than LinkedIn Ones

The severity of agent drift scales with the volume and velocity of your content output. A LinkedIn thought-leader posting once a day might be able to maintain their AI assistant manually. The corrections are few enough that they can keep a running document. But a TikTok creator posting three times a day, or a social media manager handling five platforms for a brand, is generating corrections at a rate that makes manual maintenance impossible.

I’ve run accounts where I’m scheduling 30 posts across 5 platforms in a single week. Each post requires platform-specific adjustments: the TikTok hook needs to be faster, the LinkedIn caption needs to be more professional, the Instagram story needs a different CTA. If I’m using AI to draft these, every single adjustment I make is a signal. The AI that drafts a TikTok caption with a slow, narrative hook when my audience responds to abrupt, pattern-interrupt openings is generating frustration data. The AI that keeps writing LinkedIn captions with hashtags when my client’s audience engages more with text-only posts is generating correction data.

A tool like Blume, adapted for content workflows, would cluster these signals and eventually suggest a rule: “For TikTok, always start with a pattern interrupt in the first two seconds. For LinkedIn, never include hashtags.” That’s the dream. That’s what would save me hours of manual prompt engineering every month. TikTok creators and high-volume social media operators feel this pain most acutely because their correction frequency is highest. A low-volume LinkedIn poster can survive with a decent system prompt. A daily TikTok creator cannot.

What Blume Actually Does Differently From the Incumbents

Let me be clear about what Blume is not. It is not a replacement for your AI assistant. It is not a competitor to the tools you already use. It is a sidecar application that sits next to Claude Code, Codex, and Cursor and learns from your sessions. The company describes it as a desktop app that “sits next to” these tools, which is an important architectural distinction.

The comparison points here are the memory and context features that the big AI coding tools are starting to build. Claude Code has its CLAUDE.md file, a repository of instructions that the agent reads at the start of each session. Cursor has its own rules system. Codex has similar capabilities. These are all manual systems. You write the rules, you update them when things go wrong, and you hope you remember to do so consistently.

What Blume is attempting is to automate the rule-writing process by mining your session history for signals. The maker’s response to a commenter question about thresholds is revealing: “It is quite conservative right now, either 5 occurrences in a cluster (thematically grouped, so not necessarily the exact same problem). Or 2 occurrences of something that have been identified to cause ‘pain’ (high token usage, all caps, frustration).”

This conservatism is smart. The risk with any automated rule-generation system is that it over-fits to a single correction. You had a bad day, you were frustrated with one particular output, and suddenly the system hardens that one-off frustration into a permanent rule that degrades the AI’s performance everywhere else. A commenter on the Product Hunt page, Taissa Maleh, articulated this exact failure mode: “one off feedback getting hardened into a permanent rule too early.” Blume’s threshold of five thematic occurrences or two pain-indicating occurrences is designed to prevent that over-fitting.

The other architectural choice worth noting is privacy. The product description states that it is “free to use, and uses your local harness for extracting and improving the setup. So no chats or code ever leaves your machine.” In an era where every AI tool wants to upload your data to the cloud for training, a local-first approach is a meaningful differentiator. For social media managers handling client accounts, this privacy angle is not trivial. You’re dealing with confidential brand strategy, unreleased product launches, and campaign plans that cannot be leaked. A tool that processes everything locally is significantly easier to justify to a client or a legal team.

Where the Math Breaks

The pain detection mechanism is clever, but it has a fundamental limitation that I suspect will surface quickly in practice. The system uses “high token usage, all caps, frustration” as signals for pain. These are proxies for human frustration, not direct measurements. An ALL CAPS message might mean the user is angry, or it might mean the user is emphasizing a point, or it might mean the user’s caps lock was on and they didn’t notice.

More importantly, frustration signals in text are culturally and stylistically variable. Some users are naturally expressive and type in ALL CAPS frequently. Others are reserved and would never type a frustrated message even when they’re deeply unhappy with the AI’s output. A system that relies on textual frustration signals will systematically miss the corrections of quiet users while over-weighting the corrections of expressive users.

In my experience, the most reliable signal of AI failure is not what the user types but what the user does: deleting a generated output and rewriting it from scratch, or manually editing more than 30% of the generated text. These are behavioral signals that indicate dissatisfaction without requiring the user to express it verbally. If Blume’s clustering algorithm doesn’t eventually incorporate behavioral signals alongside textual ones, it will have a blind spot for a significant segment of users.

What Creators and Social Media Teams Can Borrow From This Approach

Even if you never install Blume, the philosophy behind it is directly applicable to your content operations. The core insight is that your corrections are data. Every time you edit an AI-generated caption, every time you rewrite a hook, every time you delete a draft and start over, you are providing feedback that the system should learn from. Most of us are throwing that data away.

Here is a practical workflow I’ve started using with my own AI-assisted content production, inspired by the Blume approach. At the end of each week, I review the AI-generated content that required significant manual editing. I look for patterns: Was it always the same type of hook that needed fixing? Was it always the same platform where the AI missed the mark? Was it always the same topic area where the AI’s tone was off? I cluster these patterns mentally, then I update my system prompts and brand guidelines accordingly.

This is manual Blume. It works, but it’s inconsistent. I miss patterns. I forget corrections. I let frustration signals fade from memory. A tool that automates this clustering would be genuinely valuable, not just for coding agents but for content AI assistants.

The second lesson is about threshold-setting. Blume’s conservative approach to rule promotion—five thematic occurrences or two pain-indicating occurrences—is a useful heuristic for content operators. I’ve made the mistake of changing my entire content strategy because of one bad week or one piece of harsh feedback. The signal wasn’t strong enough to justify the change. Applying a threshold rule to your own content pivots would save you from over-reacting to noise.

The third lesson is architectural. Blume processes locally, keeping chats and code on the user’s machine. When you’re evaluating AI tools for your content operation, privacy and data control should be top-tier considerations. The tool that promises the best outputs but requires you to upload your client’s unreleased product strategy to a third-party server is not worth the risk. Local processing, or at minimum clear data-handling policies, should be non-negotiable.

The AI Content Stack Is About to Get a Memory Layer

What Blume represents, in a broader sense, is the next phase of AI tooling for knowledge workers. The first phase was generation: tools that could produce text, images, and code from prompts. The second phase was automation: tools that could schedule posts, repurpose content, and manage workflows. The third phase, which we’re just entering, is memory: tools that learn from your interactions and improve over time without requiring manual maintenance.

This is the layer that’s missing from most content AI stacks. We have Canva for design, CapCut for video editing, Buffer or Hootsuite or Later for scheduling, and a dozen AI writing tools for drafting. But none of these tools remember what you corrected last week. None of them learn from your editing patterns. None of them get better at producing content that matches your brand voice the longer you use them. They’re all stateless. Every session starts from zero.

Blume is attempting to add that memory layer for coding agents. The same concept needs to be built for content AI. Imagine a tool that watches you edit AI-generated captions in your scheduling dashboard, learns that you always remove the emoji from LinkedIn posts, and automatically pre-filters future drafts. Imagine a tool that notices you always rewrite the first line of your YouTube descriptions, learns your preferred hook structure, and applies it before you even see the draft.

That’s the future. And the team at Blume.codes is building toward it, even if they’re starting in the developer niche.

Where My Judgment Says It Falls Short

I want to be balanced here, because the enthusiasm around this launch is real and the problem is genuine, but there are limitations that potential users should understand before they invest time in the tool.

First, the product is currently focused on coding agents—Claude Code, Codex, and Cursor. The Product Hunt launch page makes no mention of content creation tools, social media management platforms, or general-purpose AI writing assistants. If you’re a social media manager who doesn’t write code, this specific tool is not for you yet. The philosophy is transferable, but the implementation is not.

Second, the product is early. The launch page indicates it’s free to use, which suggests the team is in a growth and feedback phase rather than a mature monetization phase. The maker acknowledges they “will work a whole lot on improvements going forward.” Early adopters should expect rough edges, missing features, and potentially significant changes to the product’s direction based on user feedback.

Third, the pain-detection mechanism is text-based and proxy-driven. As I discussed earlier, ALL CAPS messages and high token usage are imperfect signals for human frustration. The system may miss quiet dissatisfaction or over-weight expressive frustration. This is a solvable problem, but it’s not clear if the current version has solved it.

Fourth, the threshold for rule promotion—five thematic occurrences or two pain-indicating occurrences—is conservative by design, but it may be too conservative for fast-moving projects. If you’re running a rapid content sprint where you’re producing 50 posts in a week, waiting for five thematic occurrences of the same correction means the AI will make the same mistake five times before the system learns. In a high-volume environment, that’s five wasted hours.

Fifth, there’s an open question about how Blume handles conflicting signals. What happens when the user corrects the AI in one direction on Monday and then corrects it in the opposite direction on Wednesday? Does the system recognize the contradiction, or does it just accumulate both corrections and create an incoherent rule set? The launch page doesn’t address this scenario, and it’s a critical one for real-world usage.

Finally, the product is desktop-only, supporting macOS, Linux, and Windows. There’s no mention of a cloud version, a team version, or enterprise features. For solo operators and small teams, this is fine. For larger social media teams where multiple people are working with the same AI assistant, the lack of shared context and collaborative features could be a limitation.

What I’d Watch and Test Next

If you’re a creator or social media operator who sees the potential in this approach, here are concrete steps you can take this week.

First, audit your own correction patterns. For the next seven days, every time you manually edit an AI-generated piece of content, note what you changed and why. At the end of the week, cluster those changes thematically. I’d bet you’ll find that 80% of your corrections fall into a handful of recurring categories. Those categories are your agent drift points. Fix them in your system prompts and you’ll see an immediate improvement in output quality.

Second, if you write any code at all—even simple scripts for automation or data processing—install Blume and test it with your coding agent. The product page says it’s free, and the local processing means you’re not exposing your data. Run it for a week and see if the rule suggestions it generates align with the corrections you know you’ve been making. This will give you a sense of whether the clustering algorithm is picking up the right signals.

Third, watch the Blume team’s roadmap. If they expand beyond coding agents into general-purpose AI assistance, or if they build integrations with content creation tools, that’s when this becomes directly relevant to your social media workflow. The underlying technology—extracting human intent from sessions and using it to improve system context—is exactly what the content AI stack needs.

Fourth, start demanding memory features from your existing AI tools. When you’re evaluating a new AI writing assistant or a scheduling platform with AI capabilities, ask about context persistence. Ask whether the tool learns from your corrections across sessions. Ask whether your brand voice guidelines are applied consistently or if each session starts from scratch. The more we demand these features, the faster the tools will build them.

Fifth, experiment with your own version of the Blume approach for content. Create a “corrections log” document where you paste every significant correction you make to an AI assistant. At the end of each month, review the log and update your master system prompt. This is manual, but it will train you to notice patterns you’re currently missing. When a tool like Blume does arrive for content workflows, you’ll already have the discipline and the data to use it effectively.

The creator economy is entering a phase where AI assistance is table stakes. The differentiator won’t be which tool you use, but how well you manage the context that makes those tools effective. Blume’s approach—mining session data for human intent and turning it into durable improvements—is the right answer to a question that’s been bothering me for months. Even if this specific product doesn’t end up being the one that solves it for content creators, the problem it’s addressing is universal, and the solution pattern is transferable. That’s worth paying attention to.

Ready to Create Your Own?

Join thousands of brands creating high-performing video ads with FLOWNIB. No editing skills required.

Start Creating for Free