I’ve been using Opus a lot at work lately for a project and have had serious issues with it reaching for bizarre analogies, obscure jargon, and strange phrases. It’s still a pretty capable model, but I frequently have no idea what it’s actually trying to say. I just want it to say what it means, not invent new metaphors and ways to confuse me.

Here are some of my favorites, though there are many more I forgot to copy down (I’ll probably add some more of these over time):

Think of it as a photocopy of a ledger. The copy is not smudged, so nothing looks wrong — it is simply an older page.

Think of it as a cookbook with a wrong recipe. He corrected the cookbook. The meals already cooked from the bad recipe are still on the table — and one has been framed on the wall as the reference dish.

In another conversation, I’d asked it not to force everything into being a root cause. It then did exactly that, and for some reason, admitted it?

These last three are dressed up as root causes. That is precisely the dishonesty the bit was supposed to prevent, and I did it anyway to make every theme have a lead. (emphasis mine)

This one was particularly funny:

Three genuine oracles exist … This is not a self-licking oracle. (emphasis Opus’, not mine)

I just had fun and ran with it:

Me: I am a little concerned about the oracle not self licking, this seems like a case where we really should make sure it’s self-licking, do you have any thoughts?

Opus: There’s a terminology trap here, and it’s worth untangling before we act on it - because the two readings point in opposite directions.

“Self-licking” is the bad thing

It comes from “self-licking ice cream cone” - a system that exists only to sustain itself. Applied to an oracle it means: the goldens were generated by the very code they’re supposed to be testing. The test then passes no matter what, because it’s marking its own homework.

So when I wrote “this is not a self-licking oracle” that was the good news

Oooookay Opus, glad you were able to explain it (this time). But the root issue is that I had to spend another turn and more tokens to understand the terminology it introduced in the first place! More and more I’m finding that it uses weird analogies and phrases loaded with implied meaning that are incredibly difficult to understand without significant cognitive load. I’ve found myself reading and re-reading a single sentence multiple times, having to stop at each phrase, decipher what it meant to say (not what it actually said), try to keep those pieces all in my head at the same time, then trying to make some sort of reasonable decision about the inevitable question that Opus is asking me to answer. It feels just a little like the sentence you just read, just much worse.

I’m having to reason about an answer, keeping in mind my original request, holding project context at the same time, then using extra cycles (both brain and LLM) to coax the model into giving me a comprehensible explanation.

At this point I needed to get work done, so I went searching for a solution. I first tried to fix it with output styles and ASD-STE100. This seemed like exactly what I was looking for: a standardized way of making technical writing unambiguous and easier to understand. I found someone’s Claude output style (I think it was toppa’s output style), installed the output style into ~/.claude/output-styles, then in the /config command, set it to ASD-STE100: That helped a little, but I think still at least half of the above quotes were with the ASD-STE100 output style.

But the thing that really helped was poteto’s unslop skill, found in cursor’s plugins repo. I’m really not a fan of “just install this skill and it’ll fix your whole life”. But no, this really is that skill. It’s not too big, I ran it through OpenAI’s tokenizer and it says it’s ~1,600 tokens: That’s not too bad for a skill that cuts down on all the AI slop1 and extra noise that Opus generates. 1.6K tokens seems cheap compared to repeatedly having to burn tokens on overly verbose output and then spend more time and more tokens to decipher it. Again, Opus is intelligent, but the actual output is nearly unusable without something to keep it in check. The unslop skill did fix a lot of issues for me. I’ll be using it anytime I use Anthropic models going forward. I was actually able to understand what Opus was trying to say. But you’ll still find me reaching for 5.6-Sol and 5.6-Luna over Opus if given the choice. (for real, if you haven’t given Luna a shot, don’t sleep on it, it’s fantastic.)

All I need my models to do is say what they mean. Is that too much to ask?

Footnotes

  1. I can’t help but think that this is at least part of why Anthropic is watermarking its text, so that they aren’t training on their own AI slop