Optics vs DRAM, AI Benchmaxing, Stock Picks
This week's finds
AI Benchmaxing & The Bull Case For Frontier AI
We’re always amazed how many market commentators look at these rankings and treat them like they’re the gospel:
Currently, we do all our coding with Claude inside VS Code—the popular IDE from Microsoft—and occasionally, we try out some new models for evaluation. So, one month ago, when GLM was all the rage on social media, we once again decided to do our own test to see if we can actually use this extremely cheap model.
So, we took a task which Claude was doing for us and then we sent it to three additional APIs—GLM 5.2 Max in Openrouter, the most advanced Nemotron model from Nvidia in Openrouter, and Gemini 3.1 Pro in GCP’s API. Then, we let ChatGPT evaluate the output as a neutral arbiter—Claude Opus scored 8.5/10, Gemini Pro scored 8.0, Nemotron scored 6.5, and GLM crashed so it got a score of 0/10.
This is basically the reason why AI spend—where consumers and enterprises are actually spending their dollars—is currently close to a duopoly between Anthropic and OpenAI:
Palantir made a similar conclusion, and noted to treat all the benchmaxing which is happening with a healthy dose of skepticism, this is the company’s CTO:
“We’re excited for a new era, not one of benchmaxing but of benchmaking. The assumption that the frontier is actually the best performing is just not borne out in practice. Within 24 hours of bringing Nemotron Ultra into our stacks, we found 5 production tasks where a standard Nemotron Ultra model without post-training beat frontier models. That shouldn’t even be possible. The benchmarks are right for what the benchmark is measuring, but those are not the tasks my customers were trying to solve. And so, moving this to an empirical basis is, is how we’re going to accelerate the realization of tokens to real economic value.
This underscores that a handful of common benchmarks can be gamed. That era, benchmaxing has ended. We’ve been beholden to a small number of benchmarks that people have been designing models to and then releasing models saying like, “look how well it does on this benchmark”. But the benchmark has actually almost nothing to do with your business. So how do you figure out how to make the benchmark that represents your reality, what you’re trying to succeed at, what you’re trying to get better at and then see what model makes sense?”
So, we haven’t been able to replicate Palantir’s results with Nemotron far outperforming Claude, but naturally, Palantir is also talking their own book. They want open source models to win so that new solutions from Anthropic and OpenAI won’t be a threat to their platform.
We are positive on open-source models taking share for simple or predefined workloads—e.g. an AI model analyzing a customer invoice, classifying which comments are appropriate for a social media app, analyzing images, customer support, etc.—however, for the really advanced engineering work such as writing large code bases, this will continue to be a consolidated market in our view. The reason is simple: only the largest players will have the capital to purchase and create the largest data sets—including paying professionals to create highly customized and advanced data sets such as investment bankers building financial models, lawyers explaining how to argue a case or select a jury, etc.
AI is basically a scale game where both the quality and quantity of the data the model is being trained on are two of the most crucial factors when it comes to model performance. Therefore, the thesis that the model layer in the AI stack will fully commoditize will continue to be wrong. Bears on frontier models also cite this token usage chart from OpenRouter, arguing that open-source is clearly winning:
However, it doesn’t make sense to use frontier AI models via OpenRouter as you would have to pay a premium. So, everyone either connects to the Anthropic and OpenAI APIs themselves, or uses the big cloud platforms such as Amazon Bedrock. The OpenRouter chart is only useful to see directionally how fast token usage is growing, and which open-source models are being used. However, to draw any conclusions on frontier model usage, its value is basically zero.
The OpenRouter chart is a reason to take Palantir’s comments on Nemotron Ultra with a grain of salt—despite the model being fully free (Nvidia subsidizes your token consumption in return for you allowing your data to be used in training), people just pay for the other models instead.
Another flaw is that market commentators typically only look at the cost per token when assessing the economics of each model, however, one should also take into account the opportunity cost. Every time you give coding work for example to non-frontier models, the risk increases that it will do something wrong and you just lose a lot of time. A model which makes you lose hours because the code has bugs, or even worse, a bad fundamental design, has a high opportunity cost, no matter how cheap the tokens are.
Token cost is a rounding error against engineer time. A senior dev costs somewhere around $100–150 per hour, whereas an agentic day might burn $30–80 in tokens. So the frontier model only has to buy back 30 minutes a day to pay for itself entirely.
We’re on the $200 per month plan with Anthropic, and it definitely feels like a steal. The value of the plan is like having a small engineering team under your control at any time which delivers high-quality work, and extremely rapidly. We will occasionally continue to try other models—for example, we’ll try out Grok 5.0 Build when the new model releases, and the next Gemini Pro, however, we think that frontier LLMs will be a good business as there won’t be much competition. In addition, engineers also get used to working with a particular model, creating a stickiness effect like in software.
Currently, we simply code everything with Claude, and might consult Gemini Pro if Claude proposes something that doesn’t look like a great solution. Recently, in one of our apps, Claude proposed writing a very cumbersome piece of code, and when we gave the problem to Gemini, it came up with very elegant and simple code to implement the functionality. Claude agreed that Gemini’s code was much better.
Google—Taking Some Profits
As discussed above, Gemini Pro does feel like a smart model to us and we regularly consult it when Claude gets stuck, and occasionally, it comes up with very smart solutions. However, as we’ve been using both Anthropic’s and Google’s products extensively over the last 6-9 months, it does feel to us that Anthropic is massively taking the lead.
A key problem is that Gemini Pro just regularly crashes, whereas Claude is extremely stable. We can work with it on a large code base in long sessions. Claude also just gives extremely robust results—it doesn’t touch anything it isn’t supposed to touch, and just implements the changes you want.
Another key problem with Google is that Search is still 50% of revenues (and more of profits), and this business is likely going to get disrupted. At the same time, Google hasn’t released a new model in ages, and now also some of the best people in AI have just left the company.
The parts of the business we like are GCP, Youtube, and Android; but these are only around 40% of revenues combined. Given the disruption in Search, at 27x forward PE, this is an expensive name:
We’ve kept half of the position for the moment on the hope they’ll release a cool model soon, but Anthropic is just on a much faster innovation cadence. Our impression is that Anthropic is running away with this and that the gap with competition will only increase from here. Anthropic has the cash to build custom data sets, has the talent, and the compute. Google seems to be lacking in execution. We don’t want to write Google off fully, but it certainly doesn’t feel like high conviction.
Next, we will dive into optics vs DRAM, as well as stock picks in power semis, leading edge semis, software, and physical AI. We’ll detail our best finds from this week.





