Skip to content
Ventures Story How I use AI Blog Let's talk 🇮🇹 Italiano
tech

AI agents: how to read the marketing without getting played (the Hermes case)

· Nicola Giunchi
Conceptual illustration: a luminous sphere revealing itself to be an empty shell next to a small solid stone: substance against hype.

Every week a new AI agent comes out promising to change your life, and every week the copy is practically identical: more memory, more autonomy, and a new name for “we automated the thing you used to do by hand”. It’s the same arms race we watched on smartphone cameras around 2018, when every model boasted a bigger sensor and nobody explained whether the photos actually came out better. With agents we’re at the same point, and the noisiest case of the moment (Hermes Agent from Nous Research, pitched against Claude Code practically everywhere) is a perfect way to think about how to read this marketing without getting the wrong end of the stick.

What “self-improving” actually means

Self-improving, translated from marketing into plain language, means something far more modest than it sounds: the agent keeps a persistent memory of past conversations, writes itself small reusable routines after finishing a task, and builds a profile of who you are over time. That’s it. In concrete terms what happens is that the agent maintains an indexed database of its sessions which it consults before answering you, summarises old conversations, and generates ready-to-use skill files for itself. It’s useful, quite a lot so, but the word self-improving evokes an intelligence that evolves on its own, whereas what you actually have is accumulating memory plus routine automation. The difference isn’t academic, because it radically changes what you can expect: you aren’t buying a system that becomes more intelligent, you’re buying a system that remembers better.

Why the benchmarks they show you are worth little

When a vendor tells you their agent is 40% faster, there’s only one first question to ask: who took the measurement? With the Hermes case something even more instructive happens. That “40% faster” circulates widely, but it doesn’t trace back to any official measurement from Nous Research: it surfaces from third-party content pointing to a supposed independent benchmark of which no verifiable trace can be found. When a number doesn’t even have a known parent, you already have your answer. The figure that went viral (“it won 14 tasks out of 18 against the competition”) comes instead from one blogger’s test on eighteen prompts, and the four it lost, it lost on pure coding. Neither one is an independent, reproducible benchmark, and treating them as such is the fastest way to build yourself the wrong expectations. Then there’s the most-cited metric of all and also the most useless, the number of GitHub stars. Within a few weeks of launch the same project went from zero to over a hundred thousand stars, and kept climbing visibly: a different number every time you check, depending on who’s counting and when. Stars measure the enthusiasm of the moment and the timing of the launch, not whether that tool will make you work better on Monday morning.

The right question isn’t the one the marketing suggests

Marketing leads you to ask “is this tool better than nothing?”, and the answer is obviously yes, because anything is better than nothing. The question that actually counts is a different one: is the new tool better than the system you’ve already built by hand? Anyone competent almost always has their own stack, made of pieces held together with duct tape but working, and the new tool rarely takes you from zero to one. At best it takes you from eighty to ninety, and you pay for those ten points in setup, maintenance and context to be rebuilt from scratch. There’s a second way this framing misleads you, which is the habit of hooking the new tool onto one you already know. Hermes is presented more or less everywhere as being “for people who use Claude Code professionally”, but it isn’t a Claude Code plugin, it’s a completely different category (an always-on generalist operator versus a specialist coding agent). Confusing the categories makes you buy the wrong thing for the wrong reason, convinced you’re adding a piece to the puzzle when in fact you’re buying a new puzzle.

What it actually costs

The claim “it runs on a five-dollar VPS” is true and misleading at the same time. It’s true that the server costs almost nothing, but the real economic advantage only materialises if you actively route tasks towards cheaper models, deciding case by case what deserves the premium model and what can go on a poor one, and that is operational work someone has to do continuously (in all likelihood, you). Add API key management, the security of an agent that executes commands on the machine and runs with nobody watching it, and you realise the real bill isn’t the server invoice, it’s your time. And there’s a security detail the enthusiastic copy tends to skip: being a young and very fast-moving project, Hermes already has several publicly disclosed vulnerabilities: a couple of CVEs on authentication and one on DNS rebinding, all fixed in later versions. That isn’t a scandal, it’s normal for software this new. But that’s precisely the point: “no known holes” is never the same as “no holes”, and entrusting data that matters to an always-on agent that executes commands on your machine and has freshly patched flaws is a choice that asks for oversight, not enthusiasm.

How to decide for real, in a week

The honest way to close any of these questions isn’t reading yet another comparison online (this one included), it’s building yourself a small, measurable experiment. Take a single recurring workflow you already do, run it for a week on the new tool while keeping it in parallel with your current setup, and measure only two numbers: the minutes you save once the system is up to speed, and the hours you spent configuring it. The ratio between those two numbers gives you the green light or the stop better than any review, because it’s calculated on your work and not on that of a blogger chasing clicks. If the experiment doesn’t hand you a clear net advantage, the answer is no, and that’s perfectly fine, because having checked is worth more than having hoped.

And here comes the uncomfortable part, the one marketing will never tell you because it has nothing to sell on top of it. In the overwhelming majority of cases the bottleneck in your productivity isn’t the tool you’re missing, it’s you: the focus you don’t defend, the things you don’t delegate, the noes you don’t say, the decisions you keep postponing. The new tool cures one very specific problem and none of the others, and the temptation to spend the afternoon configuring it instead of doing the real work is one I know well myself, because it has the look of productivity while being only motion, not progress. So before asking yourself whether the latest agent will change your life, stop for a second and ask the question that really counts: what is slowing you down, right now, actually? If the honest answer isn’t “this exact tool is what I’m missing”, then you already have your answer. And maybe that afternoon can be spent a great deal better.

Nicola Giunchi

Nicola Giunchi

Serial entrepreneur, investor, writer. Founded 8+ companies in 20 years.

Frequently Asked Questions

Is Hermes Agent better than Claude Code?

They're two different categories, not two versions of the same thing. Claude Code is a specialist coding agent that lives in the terminal, Hermes is an always-on generalist operator that lives on a server. Most serious users run both, each for what it does best. The word better depends on what you need, not on which one sounds more impressive.

Does self-improving mean the AI trains itself?

No. It means it keeps a persistent memory, writes itself small reusable routines and builds a profile of the user over time. The underlying model doesn't change: it's memory plus automation, not an AI rewriting itself.

Is it worth switching tools every time a new one comes out?

Almost never. If you're competent you already have a stack that works, and the new tool rarely gives you a real jump. More often it gives you a small improvement at a high cost in setup and maintenance. Switch only when a measured experiment proves it to you.

How much do GitHub stars matter?

Not much. Stars measure hype and timing, not how useful a tool actually is for your work. The same project can show wildly different numbers weeks apart, depending on who's counting and when.

AIAI agentsClaude Codemarketingproductivity