The Token Limit Nobody Talks About
Every AI model has a context window.
It's a hard limit. A finite number of tokens—words, symbols, concepts, and relationships—it can actively reason about at one time. Stay within that window and the model is fast, accurate, and surprisingly intelligent. Push beyond it and performance starts to degrade. It forgets earlier information, loses context, makes strange assumptions, or simply gives up.
It's not a bug. It's physics.
The bigger the context window, the more complexity the model can juggle before it begins to struggle.
The funny thing is... Nobody tells you that running a business works almost exactly the same way. People have context windows too.
We don't measure ours in tokens. We measure them in meetings, decisions, interruptions, Slack messages, emails, production incidents, people problems, budget reviews, project updates, and the twenty-three conversations we're trying to remember while someone asks, "Hey, do you have five minutes?"
(For the record, nobody has ever meant five minutes.)
Unlike AI, though, we don't have a dashboard telling us we're at 98% utilization.
We just notice that we reread the same email three times. We walk into a room and forget why. We stare at a Jira board wondering what we were supposed to be working on. We open Teams to answer one message and somehow end up in four conversations that weren't even related to the original problem.
By 2 PM, we're mentally exhausted.
Not because we've done physical work.
Because we've run out of tokens.
My Official Title is CTO...
...but my subtitle is "Lives in Chaos."
Honestly, I should probably have that printed on a business card.
If you've ever worked with me, you've probably seen it.
One minute I'm discussing enterprise architecture. The next I'm talking about networking. Then cybersecurity. Then supply chain. Then budgeting. Then AI. Then someone walks in asking why a printer in Building 7 stopped working.
Five minutes later I'm reviewing Kubernetes deployments while answering questions about manufacturing systems and trying to remember whether I already approved that capital request.
It's controlled chaos... most of the time.
The reality is that executives, especially technology leaders, become professional context switchers. Our job is to absorb complexity from every direction without letting the organization feel it.
We're essentially acting as a giant context buffer for everyone else.
And that's where things get dangerous.
Token Maxxing: The Superpower... and the Trap
Over the years, I've noticed a pattern—not just in myself, but in software architects, founders, engineers, and just about every technical executive I've respected. It seems especially common in people whose brains are wired a little differently. Whether that's ADHD, hyperfocus, or simply years of training ourselves to solve impossible problems, I don't know. But the pattern is unmistakable, and I've started calling it Token Maxxing.
When the conditions are right—a difficult problem, meaningful work, high stakes, and a deadline—our brains become incredibly efficient. Information flows faster, patterns emerge almost instantly, and connections appear between systems that seemed completely unrelated just minutes before. Problems that have been stuck for weeks suddenly become obvious because you're no longer seeing ten separate issues; you're seeing the one underlying cause tying them all together. What looks overwhelming to everyone else starts feeling almost... fun.
It's one of my favorite states to be in. Hours disappear without notice. Food becomes optional. Sleep becomes negotiable. Someone asks what time it is, and you're genuinely surprised it's already dark outside. From the outside, it looks like a superpower. People wonder how you managed to accomplish three days' worth of work before lunch, and they assume you're simply smarter or somehow operating at a different level.
In reality, that's usually not what's happening. You're just consuming tokens at an incredible rate. Your brain is operating at full capacity, processing an extraordinary amount of context every second. It feels amazing while it's happening—but every system has its limits, and eventually the bill comes due.
The Crash
Then it happens.
Not gradually, either. Almost instantly.
The code you've been staring at for six hours suddenly stops making sense. Decisions that felt obvious thirty minutes ago now require effort. You reread the same email multiple times. Someone asks you a simple question, and instead of an immediate answer, your brain responds with, "I have absolutely nothing left."
You didn't become less intelligent. You didn't lose motivation, and you certainly didn't stop caring about the work. You simply exceeded your cognitive context window.
Just like an LLM that has been pushed beyond its token limit, your brain starts dropping information to protect itself. Details slip away. Connections become harder to make. Patience gets shorter—not because people are asking bad questions, but because every new question requires loading another set of tokens into a system that's already full.
For years, I fought this. I'd stare harder at the screen, convinced that another fifteen minutes would solve the problem. It almost never did. In fact, the longer I sat there trying to force it, the worse my thinking became.
Now I've learned something surprisingly effective.
Start a new session.
Not on your computer.
In your brain.
Get up. Walk outside. Grab a coffee. Talk to someone about anything except the problem you're solving. Go for a walk around the building. Pet the dog. Watch ten minutes of birds doing bird things. Give your brain permission to unload its context window instead of trying to squeeze in one more token.
It's remarkable how often the answer shows up five minutes later.
We've all experienced it. You struggle with a problem for hours, walk away in frustration, and halfway through the walk the solution appears out of nowhere. The solution wasn't magically created during the walk. Your brain finally had enough free capacity to see it.
Over the years, I've come to believe that burnout isn't usually caused by working hard. It's caused by carrying too much unresolved context for too long.
There's an important difference between those two things.
Hard work is tiring, but it can also be energizing when you're making progress. Carrying dozens of unresolved problems, switching between twenty priorities, and trying to remember every open decision is something entirely different. That invisible cognitive load slowly erodes your ability to think clearly.
Leadership is Token Management
This realization completely changed the way I think about leadership.
As a CTO, my job isn't simply making technology decisions. It's reducing unnecessary cognitive load across the organization. Every unclear priority, every unnecessary meeting, every process with fuzzy ownership, and every decision that gets revisited over and over again consumes mental bandwidth that could have been spent solving real business problems.
The best leaders I've worked with all seem to understand this instinctively. They create clarity before they create urgency. They simplify before they optimize. They remove friction instead of asking people to simply push harder.
That's why documentation matters. That's why clear ownership matters. That's why good architecture matters. Those things aren't just operational improvements—they're ways of giving people back mental capacity. Every ambiguity you remove is a few more tokens your team can spend on creativity, problem solving, and innovation instead of trying to remember who owns what.
I've started looking at organizations through that lens. Some companies are constantly asking people to expand their context window, expecting them to juggle more meetings, more interruptions, more projects, and more priorities without changing the system around them. Eventually everyone starts making slower decisions, quality slips, and the organization mistakes cognitive overload for a performance problem.
The healthiest organizations do the opposite. They reduce unnecessary complexity. They automate repetitive decisions. They eliminate noise. They create systems that allow people to stay focused on meaningful work instead of constantly reloading context.
Ironically, those organizations often accomplish more while appearing calmer.
Maybe that's the real lesson from AI.
Performance isn't just about how smart the model is. It's about how much context it can effectively manage. The same is true for people.
The goal shouldn't be to cram more tokens into our brains. It should be designing organizations where fewer tokens are wasted in the first place.
Because when people spend less energy surviving complexity, they have more energy left to solve the problems that actually matter.
Think this argument fits your event? Tell me about the room — the calendar is selective.
Start a conversation