Here are some thoughts on Jev after trying it over the weekend.
It's a classifier with a ton of training. But what excites me the most is that it brings more control to the frontier, we don't have to simply use an LLM for every AI task because those are the only intelligent enough option!
How many of us already finished a prompt with "reply with only a YES or a NO" to then have some code read the result?
We saw how good symbolic AI is when we got the harnesses that could call tools. The launch of a frontier classifier like Jev means we are one step further into the symbolic aspect, at least in my understanding.
Here's an example: imagine Alice and Bob both have their own AI assistants and they want to schedule a meeting with each other
Today, with LLMs:
• Alice's assistant messages Bob's assistant, they exchange some available slots in some arbitrary format, and schedule it.
• Maybe more realistically: Alice's assistant opens Bob's calendar link and clicks around on a web interface comparing slots with Alice's.
Some engineer could look at this and think "that's very inefficient, we should have a tool to help the process!". But I think there's a reason why we don't select slots deterministically yet - there's too much nuance, we normally don't just want something that takes the first available slot.
Adding a (good) classifier to the mixture, you can have the best of all worlds:
1. LLMs communicate to find a list of slots for each person on a given format
2. Code matches the slots to find the intersection
3. The classifier picks the slot
If Alice is not a morning person and want that to be taken into account, a classifier can do it. If the classifier confidence is lower than Bob's requirements, it can escalate options to Bob for him to confirm.
It adds more structure and control, compared to LLMs, while not going all the way to the limits of determinism where we don't want it.
I'm not saying this is it's killer use-case. I'm particularly more interested in using it in safety classifiers, as using Anthropic's auto-mode doesn't feel like the best choice (specially in light of recent hacks); and in using it for sensors and controllers/gates on the software factory setting.
LLMs have been the hammer for every single nail for too long already, finally something is changing.