Things I Think I Think About AI (2026 Edition) // BRXND Dispatch vol. 128
Revisiting AI harnesses, top models like Fable, and agent-first workflows reshaping the field.
Last year I published a list of 29 things I think I think about AI, and I thought it might be good to revisit. So here’s my 2026 edition:
But first BRXND NYC is 90 days away on 11/5, and early bird tickets are currently $749. Get yours early, we will sell out.
Things I Think I Think About AI (2026 Edition)
Everyone is going to be working in harnesses (Codex, Claude Code, etc.) in the future.
I’m not sure if that harness-first future is 6, 12, 18, or 24 months out.
If your primary experience of AI is still in ChatGPT, Claude, or Copilot, you have no idea what these models are really capable of.
Don’t count out Google and Microsoft in the harness race quite yet. I still think the fundamental question is whether these model providers can build great interactive surfaces before the interactive surface providers can figure out how to properly integrate models.
Anyone who says AI can’t do useful work is lying to themselves or to you.
Fable is the best model on the market, and I don’t think it’s close.
Fable is the first model that feels smart enough that you shouldn’t waste its time having it write simple code. Its best use is as a planner.
There are only three real vectors to judge models: Raw intelligence (frontier foundation models), cost-per-intelligence (value models), and tokens-per-second (there’s some minimum bar for intelligence here, but it’s low), everything else is competing for scraps. The frontier foundation models are where we all pay the most attention (Fable, Opus, Sol, etc.), but I think the most interesting/dynamic space right now is cost-per-intelligence, where you have Gemini Flash-Lite, 5.6 Luna, and open source models like DeepSeek Flash v4. These are workhorse models and play a huge role in almost any large-scale AI project.
I still think the hype around local models is far from the reality. The only thing I find myself using them for at all is a tiny bit of local search/embeddings stuff and almost nothing else. The good open source models are too big to run locally, and everything is just too slow to make it worth it (before you ask, I have an M5 Max with 128 GB RAM and a DGX Spark in my home rack).
Codex Desktop (now ChatGPT) is awesome and lowered the barrier to entry on harnesses in a really powerful way. I don’t understand why everyone isn’t using it all the time.
The tokenmaxxing conversation is stupid: the vast majority of companies should be spending way more on AI, not way less.
The best way to learn AI is to buy a $200 Claude or OpenAI plan and try to exhaust your limits.
There are a lot of companies that are trying to optimize their AI spend before they have any. That is a dumb mistake that will come back to bite them when, in a year, they still have no one using AI.
While token leaderboards are a blunt object, generally tracking adoption and making sure teams are using the technology is a good thing to do. Obviously it can create bad incentives, but the more damaging incentive is individuals doing nothing.
The number one challenge with using AI to write code isn’t buggy code, it’s misalignment with vision and architecture. These models are already just as good (if not better) than your average mid-level engineer, but they’re terrible at deciding what’s valuable to work on and staying in line with the technical architecture that’s been laid out without quite a bit of coaxing.
There are two primary modes for using AI: divergent and convergent. The latter is for problems where you know what you want, and you should be able to just run it through some kind of system and let the AI take care of it. If it does a bad job, it’s almost always a failure in how you designed the system, not a failure of the AI. The former is how you use these tools to think through new problems and ideas, and it’s a much less well-explored space. Here you don’t want to give over judgment to the model, you want your taste and ideas incorporated. This is one of the most interesting problems to solve in AI over the next twelve months, and we’re starting to see little signals of it in Claude Code and Codex (live artifacts).
The build vs. buy decision inside enterprises is at a very interesting moment. In most cases, I think it’s fully flipped to the point where build is the incumbent and buy has to make a case rather than vice versa.
If you’re not thinking about a company/brand brain, it’s a good time to start. Building tooling for the harness-first workforce to have access to all the systems and data it/they need is critical.
To that end, we’ll be talking more and more about agent experience (AX) and what it looks like to truly design agent-first. One of the most interesting questions in this realm is where AX drives you to make different decisions than UX/developer experience (DX) would. One obvious place is defaults. The rule of thumb we’ve set is to default these systems to human-readable output because the agents are much better at navigating complicated CLI flags and options.
Memory in chat/harnesses still doesn’t really work yet out of the box, and I’m not positive it ever will.
A lot of the productivity gains we’re seeing from AI are actually really smart people getting much better at using the models.
Evals still aren’t trustworthy. I frequently find that a model that has lower eval scores does better in real-world work than the maxed-out one.
I continue to be surprised that the new models seem to be worse writers, not better. I suspect there’s a mix of reasons for this, including the move to code/verifiability. But fundamentally I’m beginning to believe it’s because writing is fundamentally anti-consensus: if everyone is saying the same thing, it is no longer a good or interesting thing to say.
There’s a new model of work developing where you start by doing it with the AI, harden it into a skill, share the skill out to others, and then, when it’s sufficiently tested, figure out how to turn it into software. This seems like a pretty durable pattern.
In the past, the answer to “should we rewrite this from scratch?” was almost always no. That’s no longer the case.
While the labs train the models, much of the most interesting work is still happening at the point of execution where regular users find new capabilities by pointing them at their real work.
And a few from the original that I still think fundamentally belong:
27. We will continue to be surprised to find out that, without any specialized training, transformer-based models can solve many problems that specialized models were built to solve.
People are the bottleneck in enterprise AI adoption, and it will stay that way for longer than anyone expects.
If models didn’t progress from what we have today, we’d still have 10 years of runway to integrate their value into corporations and society.
When it comes to AI, don’t trust anyone who sounds too confident.
The value of using AI isn’t that it gives you great foresight into the future of how AI will evolve; it’s that it gives you an uncanny ability to sniff out everyone else’s BS.
AI is underhyped.
And finally:
This whole field is so fun. I can’t believe I get to do this for a living.
Editor’s note from Mike: We would love to hear any stream of consciousness thoughts about AI on all of your minds. Please feel free to build on the discussion in the comments or email me any takes at mike@brxnd.ai and we’ll incorporate into future editions of BRXND. Thanks, as always for reading.



