Small Models Have Arrived

For the past few weeks, I've been playing with gpt-5.6-luna. It is shockingly capable, fast, and smart. I regularly see it do ~100 tps, and rip around my codebase, email, and knowledge base.

Of course, the biggest thing with luna is the cost. I've tried running some fairly complicated research threads, and it's pretty tough to run up a large bill. Even having it search across thousands of emails, I end up with an API cost in the tens of cents.

With GLM 5.3, we even have a new option at the Pareto frontier.

When doing coding work, I almost always reach for the most expensive and capable models (Fable 5, 5.6 Sol). So it's been easy to miss the progress the small fast models have made.


One thing a few investors I've talked with have mentioned: "It's weird we're not seeing more consumer AI companies. Why is that?"

There's a straightforward answer: token costs.

In the times before AI, the playbook for big consumer apps looked like this...

  • create some sort of compelling website which is fairly cheap to run
  • attract a bunch of users (typically with some virality)
  • raise money, scale to more users
  • create an ads marketplace

This roughly describes most of the big consumer companies (Google, Facebook, Snapchat, etc.).1

But what if you want to add AI to your product? Well, now you have some real inference costs on every request! Suddenly the amount of capital required increases dramatically.

A pet eval of mine is to build a daily news site, personalized to me:

research @calvinfo on the internet. figure out what news they might like. build a micro-site with today's top stories, personalized for them. search hn, reddit, twitter, etc.

With the previous generation of models (Sonnet class), you'd spend ~$1 to get anywhere. Charging $30/mo is untenable for a consumer app. There's obviously a lot we can optimize here, but if you're charging what the WSJ or The Economist charges, you'd better be delivering similar value.

But looking at luna, the results are pretty decent, and the average cost is ~$0.10. Now we're talking!


Where I think this gets even more interesting is in the world of business.

My Segment co-founder Peter and I were recently comparing notes on a hike. Across his various startups, Peter has seen two kinds of work:

  1. the "IQ 180" work. some mad scientist genius type comes up with some crazy solution you've never thought of.
  2. the "token spewer" work. being ultra responsive, pushing the ball forward across dozens of different fronts.

Peter runs multiple companies. Beyond Segment, he's raised $100m+ for Charm Industrial, and just recently closed a Series A for Revoy. He's incredibly organized and efficient with his time.

And yet, Peter mentioned that ~95% of the work he does falls into bucket 2. It's hopping on calls. Nudging people. Blocking and tackling.

To be clear, Peter says his companies would be dead-in-the-water today without an IQ 180 technical mind solving the deep problems. Just that most of his work falls in bucket 2.2

I think demand for "frontier-level" models is going to keep compounding. Especially for fields that require novel breakthroughs or discovery (engineering, hard science, model training).

But I also think the demand for "fast/cheap/good-enough" models is just about to take off.

Think of the people you interact with on a daily basis: coworkers, vendors, and customers. Nine times out of ten, you want someone who is super responsive, and just handles things for you. Most of the "human tokens" at companies today are spent this way — hiring skews heavily toward the fast/cheap/good-enough archetype.

There's a lot of work that needs to happen to make fast/cheap/good-enough models a reality for business. New harnesses, prompt injection safety, roles, and permissions. But I'm confident we'll figure that out.

If you're also experimenting with making small models useful, please drop me a line.

Footnotes

  1. Amazon and Netflix are the notable exceptions

  2. Peter is also being modest here. He's sharp as a tack.