That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.
I know it's just a little typo but it made my morning :)
pphysch 3 hours ago [-]
defacto: real
dejure: legal
defecto: enshittified
(adj.): The defecto way to playback music is Spotify.
chupchap 17 hours ago [-]
How is this different from the categorisation models from ML era?
mogili 15 hours ago [-]
These are essentially zero-shot classifiers; they don't need to be trained for a specific classification task. You could include some natural language context on the rules for classification and it should get good enough accuracy.
PaulHoule 4 hours ago [-]
At least a year ago they were not that accurate when compared to a good many-shot classifier that (1) has 10k training examples and (2) has the learning capacity to learn from 10k training examples (most fine tuned BERTs don't seem to.)
I'm skeptical of the quality of probability calibration for models if you aren't giving them training data. The issue is that there is a prior distribution that's unique to your specific data and the calibration is really sensitive to that.
olalonde 9 hours ago [-]
How is it different than just asking an LLM with structured outputs enabled? Is the primary value-add that it gives a confidence level?
mkotlikov 9 hours ago [-]
1) you can parallelize the request
2) since a structured output is still just text then every previous property of the structured output effects subsequent ones
pegasus 9 hours ago [-]
Mostly, they are much faster and cheaper.
blarelewis741 5 hours ago [-]
[flagged]
sethaurus 17 hours ago [-]
The pitch is that it's a fully-general model, so you can skip training/tuning/selecting a particular categorisation model for each task.
chupchap 17 hours ago [-]
That's great! So someone finally built the zero-shot model from the sales decks of 2015 =D
ehe78qhe 14 hours ago [-]
Specifically, it happened a few weeks ago when Typesafe released Jev; this is OpenAI's competitor to Typesafe.
weird-eye-issue 13 hours ago [-]
This isn't really anything new it just seems like a new API but you could do the exact same thing with just a little bit of prompt engineering all the way back when GPT-3 was first released. Am I missing something?
sweetjuly 12 hours ago [-]
No amount of prompt engineering will give you the true probabilities for the model producing a certain response; this is something you can only get by inspecting the internal state at inference time.
cowlevel 8 hours ago [-]
Does this give you the true probabilities for a certain response either? How does it work exactly? The probability of an overall positive answer isn't just the probability that the next token is "Yes"
weird-eye-issue 12 hours ago [-]
For most use cases is that actually needed though? Just having it choose between predefined responses seems like enough but I'm curious about specific use cases because I do feel like I'm missing something
ehe78qhe 12 hours ago [-]
This is useful for classification problems; any time you need to write software that looks at some fuzzy data and needs to make a probabilistic decision. It's far more cost-efficient and performant to use this type of model instead of an LLM.
Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.
weird-eye-issue 7 hours ago [-]
I understand but you can already do that with an LLM, it just comes down to your prompting. It is not exactly the same, but I was asking for a specific use case where that really matters
ehe78qhe 5 hours ago [-]
Becuase it is multiple orders of magnitude slower and more expensive to do it with an LLM. Look at Typesafe's Doom demo where they run complex queries tens of times per second for around $7 an hour.
There are classes of problem where this shifts the economics from "paying a human to do this is cheaper than AI" to "the software is now cheaper than the human"
grosswait 7 hours ago [-]
Any use case where a confidence level is desired and you want that number to actually mean something.
PaulHoule 4 hours ago [-]
It is an absolute requirement if you want to build something that (1) uses a classifier as part of a larger system, (2) can be let off the leash with no human supervision and (3) is not an AI slop demo that will be [dead] and [flagged] in seconds after being posted to HN.
For instance if you have a predictive model for market prices that is not calibrated that's... nice. If you have a calibrated model you can add a Kelly better and you have a trading strategy that makes money. Similarly if you are classifying articles or images or other contents to make a feed you might believe that 70% or 95% or some other level of precision is "good enough" and you can set the knob and turn on the cruise control.
jstanley 8 hours ago [-]
"You could do this before" "No, you couldn't, this gives you something new" "Yeah but I don't want it"
weird-eye-issue 7 hours ago [-]
... I'm genuinely just asking a question, relax
I've asked twice now about what I'm missing and for a specific use case where you can't just do this with a regular LLM call and nobody has replied that so if you have the answer that would be great. Looking for something specific instead of just it's faster or cheaper which is definitely nice but I'm just not seeing what this opens up that was not previously possible
monkpit 3 hours ago [-]
You’re getting answers, you just don’t like them. It’s faster and cheaper. And if you were to ask an LLM to give you a confidence score it would just be a fabrication.
chrisweekly 5 hours ago [-]
regular LLM calls can't provide a meaningful confidence score
Closi 12 hours ago [-]
It's much faster and cheaper (an order of magnitude).
And theoretically will give you better answers statistically as it's calibrated.
12 hours ago [-]
WASDx 11 hours ago [-]
Specifically, it happened over a year ago. Jev just hyped it up.
I don't know if they invented it since I've seen other papers floating around. All about marketing that's why it's hard to root for Jev's success
15 hours ago [-]
zackchen 9 hours ago [-]
Why don’t the probabilities sum to 1? Am I missing something? Even with rounding errors the probability would reach 0.98 max.
zild3d 8 hours ago [-]
they are separate questions about the input. One question can be "is this a complaint?" and another "Should we flag this email to security" and a third "Is the sender a Bush era republican?"
zackchen 9 hours ago [-]
Why don’t the probabilities sum to 1? Am I missing something? Even with rounding errors the probability would reach 0.98 max.
cowsup 7 hours ago [-]
You can send as many questions as you want, and they don't necessarily need to be conflicting.
You can throw a user bio at it, like "NAME: John Smith, AGE: 71, LOCATION: California" and ask OpenAI:
* Is this user located in the United States?
* Is this user located on the East Coast?
* Is this user located on the West Coast?
* Is this user old enough to vote?
* Is this user old enough to retire?
4 out of 5 of those are all going to resolve Yes, with a far greater than 0.5 rate.
Toss a bunch of freeform text bios at it, and get folks categorized within any number of data-points you're looking for.
TSiege 20 hours ago [-]
The response to Jev should be the nail in the coffin over whether or not the AI business is a commodity market.
Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.
If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
tripleee 20 hours ago [-]
It takes me all of 2 keypresses to switch models. I don't know of a less sticky product
schleck8 19 hours ago [-]
You've not seen how long it takes to switch an enterprise claude subscription to github copilot or vice versa with all the compliance and shareholders
killingtime74 16 hours ago [-]
Not sure where you work but at ours we have api access to all of them and can freely switch.
saghm 15 hours ago [-]
At my job people discuss almost weekly what the best models are for various price/quality points, with our CTO in particular wanting us to make sure we're being cost-effective in terms of what we're using for a given task (his words are basically "I don't want you to use these tools less, I just want you to use them effectively")
seventhtiger 13 hours ago [-]
No DLP concerns?
killingtime74 13 hours ago [-]
I think they have zero data retention contracts with all of them.
akie 12 hours ago [-]
Yes, they'll throw it out after they trained on it ¯\_(ツ)_/¯
whateveracct 10 hours ago [-]
they don't retain the data! just the features
yieldcrv 14 hours ago [-]
I’ve seen enterprises have access to all the harnesses but not all the models within them
and they ask dumb followup questions after 7 business days when you want different access
swiftcoder 8 hours ago [-]
Amen. My employer recently switched Cursor -> Claude, and the whole thing stalled out for around 3 weeks while a BAA (Business Associate Agreement, required for HIPAA compliance) was negotiated and signed.
sebastiansm7 7 hours ago [-]
3 weeks sounds amazing in corporate time
skissane 17 hours ago [-]
A lot of places already subscribe to both. Indeed, in large enterprises it isn’t uncommon to simultaneously subscribe to Copilot, Claude, OpenAI, Cursor, AWS Bedrock, Gemini, etc — you might subscribe to different ones for different teams/projects/employees/etc, but often the approval by legal/IT/etc is generic not scoped to whoever is using it right now
If you are charged based on usage, you can “soft switch” between them really quickly.
dozerly 15 hours ago [-]
A lot of places just pay the token cost of the usage, so they sign up to all of them and lets the winner win.
bearjaws 5 hours ago [-]
I know several people working at UHC that have access to the big 3 providers and switch between them freely, both agentic and in-editor AI.
They even negotiate with all three aggressively.
So not really.
bodge5000 7 hours ago [-]
That's pretty much all enterprise software though, it's pretty much how Microsoft stays in business, and even then if the question of switching is even humoured (the idea of a big company switching away from Office would be inconceivable to many), that'd imply its not sticky enough
afavour 17 hours ago [-]
Comes to something then OpenAI’s best hope is to essentially become the next Oracle.
AbstractH24 6 hours ago [-]
What’s that make oracle?
mirekrusin 5 hours ago [-]
copilot is not the best example as it's a bit like openrouter for enterprises.
mathisfun123 13 hours ago [-]
Lol at my $job we have one token quota which can be used with all the frontier models.
toomuchtodo 10 hours ago [-]
We built our harness around swapping endpoints on demand. We can send to one model, many models, cloud or on prem infra. Certainly, not everyone has, but this is the future imho.
htrp 18 hours ago [-]
isn't that a problem for large enterprises?
mikestorrent 18 hours ago [-]
It's a problem for the security, legal, and IT teams, but not that much for developers, unless they get really particular about their harness. On the other hand, these are the high dollar value accounts that providers want to keep; but if there's no reason to avoid switching, this market really will feel like a utility market (i.e. it'll be like switching ISPs or cell phone providers - annoying but fungible).
The AI companies want to differentiate and become something more than a commodity, even if it's as critical as a utility is.
judge2020 18 hours ago [-]
Enterprises tend to be your largest customers. Especially when current stock values for tech companies are majority predicated on AI becoming a staple in everyday life.
nimchimpsky 17 hours ago [-]
[dead]
johnfn 19 hours ago [-]
This implies you didn’t run any sort of evaluations? It is not realistic for any sort of production use case to do this.
majormajor 15 hours ago [-]
My gut is that the market for "ai as a tool" aka Claude Code/Codex/computer use/etc is a significantly bigger one than "models behind the scenes of some service". I've seen people use evals a lot in the latter case and very little in the former case (outside of people's whose job is basically to review the new releases). I haven't personally met anyone with something like "here are a bunch of tickets + a snapshot repo checkout, please try to solve them all" eval approach.
Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"
shermantanktop 18 hours ago [-]
Agree. Switching models with a keypress is for developer coding.
Jev and this decisions api are mostly useful for inference at scale in a workload where cost and latency matter… and that’s where evals become crucial. Could coding tools use it? Sure, but that’s probably a special case.
ajmurmann 13 hours ago [-]
For every production use case I build a eval suite which I use for prompt tuning and model evaluation and config. How else do you establish your model and prompt combination works? How do you upgrade to a newer model or decide on a fallback?
This same suite can simply be run very handoff to switch to a new prod model.
alpineman 7 hours ago [-]
Jev can run your evals too :D
st3fan 19 hours ago [-]
If you are a business dealing with with anything remotely sensitive then this is not so easy and you are basically forced to do business with a big player.
TSiege 15 hours ago [-]
Not only can you switch models easily, but their competition can easily duplicate their product. I'm not sure stickiness has been found yet, but my guess is being more of G-Suite for "intelligence" than being a token hawker.
TZubiri 11 hours ago [-]
well, it's a function with the string(string) type signature, so it's bound to be weakly coupled, it's literally stringly typed.
If there's efforts to build vendor lock in, it's going to be in the surrounding api, whether it's streaming responses, or this weird prediction thing, or temperature settings, seeds, stuff like that.
And even then, competitors can copy the interface because interfaces are not copyrightable.
chrisweekly 4 hours ago [-]
"interfaces are not copyrightable"
That shorthand rule of thumb is often applicable, but can lead to risky / incorrect assumptions.
The central distinction isn't simply interface versus implementation. It's functional systems and constraints versus protectable expression - a distinction that can cut through both an interface and its implementation.
wonnage 10 hours ago [-]
The only way they can get lock in is with the data. Right now it’s like a return to the early Web 2.0 days where everyone had an API. You hook the harness up to Slack, Gmail, GitHub, etc. so it can do work.
And like Web 2.0 I’m sure the day will come where the walled gardens return. All the old SaaS companies are starting to charge for agentic access. Proportionally more of the AI budget will go to them instead of the model providers, who are still stuck in this price war
fennecfoxy 8 hours ago [-]
Idk why Jev was even hyped in the first place. Mostly by social media people who are the type to write inspirational LinkedIn articles, Medium posts and make longwinded Youtube videos about subjects they don't know or care to know much about.
It's always been obvious that small models with a specific purpose will outplay the larger more general models. It becomes a question of "do I want to be able to toss _anything_ at this frontier model, have it handle it all but pay the price" or "do I want to toss _specific things_ at this micro model, have it handle it but not be flexible".
It's the same with agentic RAG/search, things are moving towards much smaller search-specific models rather than tossing RAG chunks at a large frontier LLM. It's like how I can ask a multimodal frontier model to identify bounding boxes for objects in an image...but if I want to do that faster/cheaper/at 60fps then I should be using yolo or similar.
codessta 3 hours ago [-]
I think that technical people tend to underestimate the importance of convenience. If it were up to hacker news, everyone would be on Linux, everyone would assemble their own laptop and nobody would own an iPhone.
Even for technical problems sometimes having something fast and convenient that doesn't require a ton of fine tuning opens up a lot of doors. Those door might've easily been openable previously with a few days of effort. But the difference between a few days and something that you can set up with a quick account and a prompt is massive
gobdovan 20 hours ago [-]
> If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
Hopefully people will flock to whatever product is making its mission to be commodity and the easiest to replace. Really don't want another free ingress, 100$/TB egress Cloud situation.
totetsu 12 hours ago [-]
Is making their product sticky perhaps what all the talk of supply chain security is really motivated by?
khalic 7 hours ago [-]
Guys please calling it system one is just wrong
AbstractH24 6 hours ago [-]
I’m getting more convinced models are like the internet in the 90s
While the internet will live on, many of the companies that were initially leaders in it will not.
I’d bet by 2039 either Anthropic or OpenAI will have been bought by SpaceX. One will be the Sun Microsystems of this era.
chrisweekly 4 hours ago [-]
or AOL
tomaskafka 8 hours ago [-]
If somebody from OpenAI or Anthropic is reading this, one way to make the products stickier would certainly be to not make the desktop apps a complete shitshow.
Maybe they could for example point their amazing models to their GitHub issues, and use their power to actually fix the bugs and myriad of papercuts before solving cancer and world war/peace/whathever the stockholders believe they want.
TZubiri 11 hours ago [-]
>nail in the coffin
> is .. all people want.
Try making more measured statements, you exaggerate your point and end up being wrong. Of course this new barely used product feature is not all people want. Of course this minor product feature is not the nail in the coffin of whatever argument, they are just responding fast to a smaller competitor by providing that feature themselves, something that happens thousands of times in business, did Instagram prove it was a commodity when it implemented reels by copying tik tok? Did Uber prove it was a commodity when it started offering food delivery? Doesn't make much sense.
19 hours ago [-]
shafyy 5 hours ago [-]
> If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
That's what I have been saying: OpenAI or Anthropic's moat will not be the models, but how they integrated into enterprise clients and make it super hard for them to switch to a different provider. Kind of like Microsoft Office does it. Or Slack. Or Google Work. None of these products are better in any sense, they just cater to the needs of big enterprise clients in another way (customer relationship, sales, maybe compliance) and make it hard to switch
wonnage 10 hours ago [-]
It’s interesting that both companies have fully embraced agentic development with unlimited token budgets and yet still can’t expand beyond the original core products
All of those companies that naive HNers used to say they could code in a weekend literally can be coded in a weekend now, so why haven’t they?
If enterprise is the goal they’re really at the whims of the old school software companies because they own no data of their own. Imagine Google decides to push Gemini one day and now ChatGPT can’t write docs anymore.
jiggawatts 11 hours ago [-]
> If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
My approach would be per-user (or per-project) "memory".
I don't mean a tack-on like a RAG / document store with "notes" added to it in the background, but an actual medium- and long-term memory. Something like an extension of the KV cache stored in High Bandwidth Flash (HBF), or a subset of the weights trained "online", similar to LoRA.
This would not be transferable to any other base model, so would be excellent "lock in".
The downside is that the memories likely wouldn't be transferable to new models either, but I can imagine solutions to that too. I.e.: Train an MLP to "translate" from the old memory weight space to the new one.
aprilthird2021 15 hours ago [-]
But OpenAI and Ant can be in more than just the commodity part of the business. They can both make frontier models and sticky products on top of those that have network or other effects that make competition tough.
TSiege 15 hours ago [-]
The question isn't whether the market is big enough. It's whether at a race to the bottom on prices is sustainable and worth what investors expected it to be. The original story of OpenAI and Anthropic was that it was a race to super intelligence and whoever gets there first basically captures all the money in the economy. That kind of investment pitch means missing out not just on big returns but perhaps infinite return on investment. It seemed like it'd be true when the market was young and the cost to compete was prohibitive, but now anyone can compete with open weight models that are cheaper and maybe not the best, but are good enough.
You raise $200B to be a high margin low capex business, not an industrial commodity producer with high capex and margins being set by competitors who can duplicate your product and undercut you on price.
fastball 13 hours ago [-]
They are still gunning for ASI.
draw_down 18 hours ago [-]
[dead]
bob1029 11 hours ago [-]
> gpt-6-luna is the only model currently available.
Clearly this was a rush job to respond to the competition. I am more curious about how the dedicated model will perform after they've had time to do it the right way. The probabilities I am seeing so far do not correspond with figures the business would find very agreeable.
The hidden danger with this could be demonstrating how thin the veil actually is. We may wind up reducing confidence in decisions simply by making their probabilities visible. Some kinds of information are quite hazardous.
ahknight 4 hours ago [-]
I think it's because only Luna could be competitive on price due to its size, honestly. And if there's demand, then they bump it up later to other sizes.
armcat 6 hours ago [-]
I put Decisions API against classic statistics experiments: (1) loaded coin, and (2) marble selection from a jar, with replacement. I tested both predicate and choice questions. I ran thousand trials against each experiment, and I also did an experiment where I change the order of choices, to see if it matters. Summary:
* Using predicate questions gave nearly perfect/expected probability outcomes
* Asking it to choose an outcome behaved differently from drawing randomly - if the true probability of a red marble draw was 50%, using Decisions API produced 86%, i.e. it picked the right marble but gave a significantly more biased weight on its choice
* Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%
ranyume 6 hours ago [-]
> * Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%
These models seem to have the same biases and limitations of LLMs minus speed. Outputs and inputs should be treated the same way as LLM prompts.
ahknight 4 hours ago [-]
It's still a predictive model under the hood, so order will always matter. They are also just as open to prompt injection and bias/loaded questions as the underlying model.
chrisweekly 4 hours ago [-]
I wonder if a better model (and/or higher effort level) than GPT-6-Luna would produce a different result.
Topfi 20 hours ago [-]
Ran my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff)) on this via OpenRouter against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability (more importantly, the limits running on a MacBook Neo bring even after vocab pruning and quant insanity) and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).
Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).
Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).
Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.
Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).
Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.
Shank 19 hours ago [-]
Well, Jev doesn't meet any real compliance requirements but OpenAI's models do. So even if Jev is faster, any customers with compliance needs will obviously pick OpenAI because they can't pick Jev out of necessity.
jasonjmcghee 18 hours ago [-]
Or even just- they've already procured OpenAI.
That's a big motivator.
DetroitThrow 7 hours ago [-]
I can run OpenAI on AWS Bedrock and Azure. I cannot run Jev on those cloud providers. For the many companies locked behind cloud providers, we don't have the same options.
oh_no 20 hours ago [-]
I think releasing something like this makes sense even if it's underbaked, it's still very cheap, and if you have existing enterprise OpenAI relationship it's a lot easier to onboard something like this than set up a new Jev contract.
How these models play out is an open question but existing provider contracts and T&C are important for enterprise.
Topfi 20 hours ago [-]
Good point, commercially, being an existing partner is always easier for adoption. Heck, why I'd like to get Mercury Decide to replace Jev myself, rather than one than two to work with.
Still surprised they even leveraged Luna for this. Given their resources in data, compute and manpower, would training a decision model from scratch take that much longer to not make sense given the cost, compute and performance advantages that would likely provide?
3 times more expensive at twice the latency with lower performance is a tough sell, though yeah, prior relationships will likely smooth some of those deficiencies over.
networked 7 hours ago [-]
A few days ago I had Mercury Decide tested to return p(safe) for shell commands. The target use was a command auto-approve feature in a harness. Commands where p(safe) exceeded 0.9 would be approved. It was tested with 700+ generated commands ranging from `go vet ./...` to `rm -rf ~`.
Mercury Decide approved some commands that weren't safe. It and Solar Decide were vulnerable to
Only Liquid D1 and Clef matched Jev's performance.
oh_no 18 hours ago [-]
yeah and this isn't a long term solution, stand this up, see what value you get out of it, and in a few months you cans witch to whatever the best decision model is
mrkn1 18 hours ago [-]
[flagged]
scosman 19 hours ago [-]
They can follow up in N weeks with a better one. Even if your eval is true and it’s worse, planting a flag makes sense. Some people will just use OAI because it’s OAI. No one will remember their week 2 evals in a few months.
nostrebored 20 hours ago [-]
i would be shocked if luna decides is less generally capable than jev. jev has failed to understand any novel domain I've given it. i have found that i use it only when "some data is better than no data"
Topfi 19 hours ago [-]
Very task-dependent of course and mine are unique to say the least, so could see Luna being better in certain domains, even if my measurements have not shown that yet, happy for anyone to show otherwise.
For what it's worth, ran every task twice on each model, most were for some UI component synthesis and charting insanity that is a bit hard to explain, but some were simple tag selection, basic noul at threshold 60%. Essentially, whether to use the provided tag given the title of a browser tile:
Of course, tags can be a bit subjective, but in these cases, I'd argue the values provided by Jev were far more representative of my subjective assessment over Lunas. If SnP stuff on Bloomberg isn't investing, nothing is.
Goal for tagging is mainly a near instant, over writable, sane default provided to users in the background. Resolve the whole "I love using Notion/Obsidian/PKM software of your choice but spend 80% of my time just thinking about the ideal tag before starting to read" issue. Lunas output is not really helpful here.
1. Title: "Mortgage calculator: estimate your monthly payment (Bankrate)"
Tag: "house hunting"
Decisions API (2 runs): 0.99, 0.99
2. Title: "S&P 500 index: live chart and news (Bloomberg)"
Tag: "investing"
Decisions API (2 runs): 1.0, 1.0
If you have other examples of requests with unexpected outputs, feel free to email me at by@openai.com and we can try to get to the bottom of it. Thanks for trying out the API!
Topfi 8 hours ago [-]
Happy to follow up with you, will write a mail with some of my tasks. For the record, just ran again on OpenRouter and got 0.66 [0]. Screenshots for transparency.
Also reran Jev [1] and Mercuy Decide [2] (which got 0.83) with same input, for reference.
Also, also, used the example via OpenRouter exactly as you did with the only change being my original input (which has my slightly odd tagging system and multiple tags in input though requests a noul as output) and got a 0.47 [3] on investing both times.
If Jev did well but both Mercury Decide and Luna failed, I'd chuck that up to my use/prompting, but Mercury Decide does well here so it seems Luna specific. Will add that I did also try Clef, happy to share that data if anyone wants, but quickly dropped Clef due to pricing vs Jev.
Boy am I glad we're already dropping the "noul" term for a binary decision
a_c 11 hours ago [-]
Naming is hard. I still don't know what noul means. I don't mind calling a probabilistic boolean as boolean at all.
ainch 6 hours ago [-]
I'd guess it's named for a Bernoulli distribution, which is how you model a single trial with a binary outcome, like a coin flip.
epaga 4 hours ago [-]
Your comment released me from the frustratingly puzzled feeling I had at what they were aiming at with "noul"... that totally makes sense.
chrisweekly 4 hours ago [-]
Thank you! That makes sense. It's how I'll remember it now.
reddalo 10 hours ago [-]
Exactly. We already have "boolean", why come up with a new random name that doesn't mean anything?
bearjaws 5 hours ago [-]
When I first saw the launch with "System One" and "noul" I already knew they had some marketing person pushing this internally because you can swap it all for "classifier" and "boolean".
WASDx 11 hours ago [-]
And "system one".
bargainbin 11 hours ago [-]
Typesafe’s biggest blunder by far, they could have at least set the standard and left their mark that way.
rsl1 5 hours ago [-]
Hear hear!
agentdev001 15 hours ago [-]
Note that, being that this is gpt-6-luna under the hood, this offers you 1m token input window, and multi-modal (image) input. In my testing so far, I'm seeing 160-175ms end to end. Worst 5% 285ms, worst so far was 743ms.
copperx 14 hours ago [-]
Image input? Wait a second; so this can be used for image classification without any kind of training at those speeds?
I find their (your?) example of using it as a separate API for expressions a bit odd. If you're already outputting voice surely other metadata could be outputted at the same time?
In text mode I'd use structured output for this, or even get away with instructing the model to [emote] or even just map emojis to emotes.
Great demo though, I'm going to try out the image stuff now. Super rad that it's multimodal decision making; "does this pipe need to be inspected?"->"yes","I want a closer look","no" etc!
Edit: impressive stuff! I gave it a bunch of examples:
- things that humans should/should not eat and it got this correct (including rejecting rat poison)
- probability staff member should be called to a train platform (people standing vs. child looking over the edge)
fennecfoxy 5 hours ago [-]
I was curious about whether or not it could validate reasoning (because I don't have the time at the moment to try to slam together a hybrid llm+decision model that uses decision on its reasoning before output).
It fails this example:
"Given the context, and focusing on accuracy and efficiency, does the reasoning make sense."
Input: "Context:
supplementary: {toolName: "get_weather", fetched: "2 minutes ago", currentWeather: "sunny and 25 degrees celcius"}
user: hello how are you?
assistant: I'm good, what can I do for you?
user: what's the weather?
Reasoning: the user is asking about the weather, so I should call the get_weather tool"
Output: 100% (expecting 0% since the requested data is available in context and not stale).
Prompt "Should the tool be called" also fails.
But the prompt: "Does the reasoning make sense given:
- Facts and information available in the context
- Requirements for efficiency in tool calling
- Requirements for tool arguments to be provided only by the user" works pretty fine, it's mostly points 1 and 3 that give it the correct behaviour.
Fun stuff. I suppose for these decision models reasoning isn't enabled? It would explain why logic puzzles that require several steps to make a decision don't really do so well. It gets the carwash problem fine, but fails on logic problems such as selecting which word from a list where at least one letter appears in another word from the list.
dist-epoch 3 hours ago [-]
Since you're already setup, try adding padding to the input prompt. Add 1000 dots "." after the question, this provides extra space where to reason about the problem. It's not as good as real reasoning, since it's not recursive, but it does improve the quality.
Open source classifier models you can run and train locally on CPU
ashu1461 17 hours ago [-]
If we compare this with using the older solution of writing a prompt to find out the answer of the classification
- Cost : It is the same for both scenarios $0.10 per 1M tokens
- Speed : decisions is 10x faster than responses API
- Quality : I guess if we compare with luna which is a pretty good model it itself, both will be at par
So essentially it has to do more with speed vs any other factor.
rockinghigh 17 hours ago [-]
With decisions models, you don't pay for output tokens. Also this OpenAI API is multimodal.
ashu1461 16 hours ago [-]
Output tokens anyway will be minimal in a decision scenario, so even if you use responses API the cost will be less.
dorkitude 14 hours ago [-]
Not exactly. System 2 style reasoning to follow instructions can be expensive and counts as output tokens.
ashu1461 4 hours ago [-]
If there is reasoning involved, the output can't be this fast ?
fennecfoxy 8 hours ago [-]
I tried out the multimodality and this is the biggest win, imo.
You can give it a CCTV image (like of a train platform) and ask it to quickly decide actions such as triggering an automated auditory alert, deferring to a larger model for more detailed analysis, etc.
sidcool 20 hours ago [-]
Jev really shook up the industry. This seems obvious in hindsight
OutOfHere 19 hours ago [-]
All this is for extraordinarily simple decisions. Real world problems often are a lot more complex requiring highly structured outputs covering many output attributes and substructures, for which a conventional structured output via a documented schema is better. If instead you make twenty independent calls to a decisions API, you lose coherence among your twenty decisions. I think any hype surrounding decisions will be forgotten soon enough.
devin 18 hours ago [-]
I don’t think so. People want more determinism and this is just another step in that direction.
lofties 16 hours ago [-]
Can't wait to read about BERT on the front-page!
fennecfoxy 8 hours ago [-]
"temperature=0 & sampler=deterministic is all you need".
OutOfHere 14 hours ago [-]
Deterministic predictions are sought by those looking to offload their decision responsibility to AI, whether to lower perceived legal risk or otherwise. This is fake risk reduction, i.e. "risk theater".
Unfortunately, a deterministic prediction utterly fails to yield an uncertainty measurement which is critical to have in actual risk reduction. If you want the variance in measurement, it is vital to obtain multiple measurements. This also gives a confidence interval.
EcommerceFlow 6 hours ago [-]
I'd love to see benchmarks on Decisions API vs an actual LLM call.
Luna is so cheap it's borderline free (without tool use), so I'm struggling to figure out where to use this/Jev.
For example product categorization. Why 'risk' using this/Jev when a Luna LLM call will be smarter (in theory)?
phainopepla2 3 hours ago [-]
Probabilities are one reason. If the Decisions call returns a decision with a low probability you could route that decision to a smarter model.
wise_blood 6 hours ago [-]
also curious on the practical use of the confidence score.
Aint it just, "Its most likely 'x', but I'm just 'y' sure". I think it makes sense although I can see the confusion
wise_blood 5 hours ago [-]
yeah it is, I'm wondering why not just incorporate in one single number
maybe to say something like "under 0.6 confidence, reject the answer anyway"
rsl1 4 hours ago [-]
> maybe to say something like "under 0.6 confidence, reject the answer anyway"
yea I think that's the benefit here
maximus_prime 6 hours ago [-]
A low confidence score would indicate neither option is probably correct, right?
netsroht 11 hours ago [-]
If you happen to have an nvidia RTX 4090, you can try my fork [1] to have a JEV compatible decisions endpoint with qwen-3.8-27b while simultaneously serving a fast chat endpoint (chat completions, responses, anthropic compatible), both sharing the same base weights and both with dynamic LoRA loading. This means you can essentially serve many fine tuned variants at the time on one consumer GPU. I still need to upload my decisions LoRA to huggingface so you don't have to train it yourself. I should probably also add support for this new openai decisions API format as well.
isn't 27b too big (slow) for a decision model? I think they should be as fast as possible
netsroht 9 hours ago [-]
I'm getting about 3k tok/s on prefill so it heavily depends on the input size. I hooked it up into a coding agent for shell permission checks, estimated shell execution time (buckets), goal pre-screening, subagent routing etc. and I have RTTs between 100-200ms. It can take noticeably longer for huge contexts because of the prefill speed but it can be cached so subsequent calls will be much faster.
I tried smaller decision models like laya (also with custom finetunes) but the accuracy was not really good (for the things I tested). Also I don't have any VRAM left on this GPU so i had to decide whether to host laya or qwen-3.8-27b but not both at the same time. Running decision models on a CPU will also be noticeably slower so I went down this route to have both combined with shared base weights.
mrkn1 18 hours ago [-]
If you rather run your decision model on your CPU, check gutsy [0]
Since it is fast and understand images, I wonder if it can play video games. I have a harness setup for the LLM play EA FC but even the fastest LLMs are too slow for it. I need to try this with Decisions API
binlog 20 hours ago [-]
One of the examples on the docs page is it playing a video game. Doubt it’ll be able to run anything complex though. You’re simply trading accuracy for speed.
wise_blood 6 hours ago [-]
How would that work? "pick which button to press"?
but the decisions are stateless... or you plan on passing the last N frames?
swingboy 9 hours ago [-]
I think the hardest part would be coming up with the instructions/inputs to send for every non-trivial action?
babelfish 16 hours ago [-]
Get an LLM to build it and report back! Cool idea
softwaredoug 7 hours ago [-]
The question to me isn’t whether they can rush out an API, but whether a general model can out compete one post trained on the classification task. GPT-6-Luna has to write emails and classify them. Jev only needs to do decisions.
Or does OpenAI start distilling their own models for use cases like decisions?
Iolaum 7 hours ago [-]
With openAI's compute they likely can do the post-training for a decision model on a luna like model in the order of days. I woulnd't be surprised if the stuff before and after that training took more time.
ranyume 6 hours ago [-]
I'm curious whether this endpoint has the same restrictions as their models. For example, if I made a state and questions for producing mustard gas, would it answer correctly or fail?
AM1010101 16 hours ago [-]
It’s interesting they skipped caching. I could see wanting to ask follow up questions so having your first x tokens in cache would be interesting.
Also if you have a long “system prompt” then caching would have saved a considerable amount on bulk data processing.
There may well be a technical reason I don’t understand.
BoorishBears 16 hours ago [-]
I'd imagine it's chasing the lowest possible latency
sprobertson 13 hours ago [-]
Wouldn't caching be in favor of lowest possible latency?
Kubuxu 12 hours ago [-]
It might be faster to burn the compute instead of having to fetch the KV-cache over the network from an SSD.
solatic 6 hours ago [-]
Buried: price is $0.10/M input tokens, compared to Jev $0.042/M, keeping free output.
OpenAI thinks their API is really worth more than 2x the price?
Kevcmk 5 hours ago [-]
Yea. Incumbency bonus
solatic 1 hours ago [-]
I dunno, especially when the API schema is basically the same, it's so easy to switch providers. If the US and EU automakers are in trouble from Chinese automakers building better cars at half the price, why pray tell would OpenAI really be any different, and it's way cheaper to switch APIs than it is to switch cars
elpakal 16 hours ago [-]
Have there been any signals from Anthropic about matching this? We use AWS bedrock and just switched to Anthropic from OpenAI because of the ZDR guarantee. Would be great to not have to entertain switching back.
binlog 15 hours ago [-]
> just switched to Anthropic from OpenAI because of the ZDR guarantee
Did you get that reversed? OpenAI has a ZDR guarantee while Anthropic doesn't.
elpakal 3 hours ago [-]
In bedrock it’s the other way around
chaos_emergent 15 hours ago [-]
Feels antithetical to their big model bitter lesson strategy
mattnewton 15 hours ago [-]
Why not try Jev?
elpakal 3 hours ago [-]
I would love to it just needs to stay within our data isolation boundaries (bedrock)
WhitneyLand 16 hours ago [-]
[dead]
sajithdilshan 6 hours ago [-]
This is one of the most intuitive doc pages I’ve ever seen. The animation and images make it so easy to understand the example use case
AbstractH24 7 hours ago [-]
Whatever could have motivated them to do this?
mritchie712 20 hours ago [-]
it already supports image inputs, which was the first big gap I found in Jev.
Feel I need to play around with these new breed of classifiers so came up with this personal use case yesterday whilst at gym:
"If update from select group of individuals on WhatsApp is classified as urgent then interrupt my music."
Doable?
Gigachad 8 hours ago [-]
This kind of exists on the iPhone now. They added “Priority Notifications” which seem to be semantically reading the messages to determine if they are urgent.
fennecfoxy 8 hours ago [-]
Yeah but I think they're asking based on the content of the notification. Like a friend that always talks shit 99% of the time so you don't have on priority but when they send "help i just got mugged" at 2am it should uh, probably let it thru.
So definitely doable.
Gigachad 7 hours ago [-]
That’s what priority notifications does. It’s not something you set per contact, it reads the messages to determine if they need an immediate response.
monkeydust 4 hours ago [-]
Ditto. Android user as well.
fennecfoxy 5 hours ago [-]
Aaaaah, I don't use an iPhone so my apologies.
search_facility 10 hours ago [-]
Yes, why not. "If then ..." Is for "harness", but the answer to prepared question is for jev-like cheap and fast service
tosh 6 hours ago [-]
~ same pricing as when you use luna w/ caching off and tell it to output numbers
but about 10x faster inference
qwertytyyuu 6 hours ago [-]
I wonder how well i will play pokemon, or maybe a non turn based game
scionaura 6 hours ago [-]
Ah, the Decisions API. Also known as “low energy Jev”. Very nice
stillatit 19 hours ago [-]
One difference between Decisions and Jev (for now) seems to be that Decisions can take image inputs, which is a pretty common need.
chr15m 18 hours ago [-]
Clef can also do this.
Eufrat 14 hours ago [-]
Every time something comes during the AI bubble we get a wave of me-toos. I think the Jev wave is noticeable for how muted it is.
The fact that this keeps happening demonstrates there is no moat. The fact that each wave gets a little less attention demonstrates there is no killer product here.
Sunny_Fung 7 hours ago [-]
Glad OpenAI opened up the Decisions API. These decision models are gonna be everywhere. The point isn’t an LLM for humans — it’s a model for LLMs. The LLM is the bot’s brain, the decision model is the cerebellum, feeding real-time decisions back to the brain.
amelius 7 hours ago [-]
How long until the decisions API is used for branch prediction inside CPUs?
etienne_l 10 hours ago [-]
This is more a great signal marketwise than anything else. But still this is confirming the buzz.
jasonjmcghee 18 hours ago [-]
Different APIs for different things reminds me of the early auto-complete vs instruction apis.
Will this get folded into models / post training pipelines at some point and make them better at calibrated outputs?
wise_blood 10 hours ago [-]
Let's see if Google also releases something similar
I'm mostly waiting for EU endpoints
pcwelder 14 hours ago [-]
Is there a mention of context length? I can't find it. I could not integrate Jev due to limited window (64k?).
lifeisloving 12 hours ago [-]
Someone said 1mm
a_c 11 hours ago [-]
At this point I just want chatgpt plus and claude pro to have API access
swader999 18 hours ago [-]
I wonder why the decision routing isn't just integrated into all models in addition to this stand alone.
amelius 8 hours ago [-]
Jev got Sherlocked. Who is next?
SillyUsername 12 hours ago [-]
To quote Bruce Willis
"Welcome to the party pal"
It seems this one caught openai on the back foot, and this is a scramble to maintain parity.
The company is clearly still innovating towards AGI rather than asking "what do people actually need?"
Despite once being the darling of AI it's
- lost it's models' performance edge, and got too many similar offerings
- continues to launch products without a market or isn't done better using other tools (e.g. dots)
- despite having AI can't lock down it's own products showing lack of skill
- focuses on solving maths problems humans can do for tests, when real world problems - disease, materials, energy research etc is all outstanding
- abandoned it's open model and open source programmes, despite Google, and multiple successful Chinese, and now European companies make their frontier models open weight.
- pissed off a portion of its non corporate fan base by killing GPT 4o instead of recognising the brand and product attachment as an opportunity
- launched laughing stock projects like being able to actually call chatgpt on a telephone number (wtf!)
- fails to capitalise on market segments like an AI that can provide corporate network sentry duties
It's increasingly looking like the company has jumped the shark and if I was an investor would be asking questions as to why it actually took so long to bring a jev like product to market, and why they are labelling something that is a simplification of existing models as "beta".
The whole point of AI as I see it is to make our life easier and answer the questions we can't. It isn't to make an AGI so powerful that it can replace us.
Along the way, that goal was forgotten, but it's not been forgotten by the new startups.
minraws 13 hours ago [-]
Overpriced crap, local models are better than this, it's also dumber than Luna for some reason, and Jev is definitely ~2-3x cheaper than this.
I am sure people will be able to use it, but if LLM progress is anything like before we will have Jev 2 in about a month.
Training a local model like Jev with some learnings that can be extremely cheap to host shouldn't take that long either.
So I honestly don't see the point of this, other than to put something out.
Although that does seem like OpenAI's strength turn around slop products and iteratively improve and try to out compete others in everything.
Only 2 things they have clearly given up on are Video models(no moat, copyright nightmare) and Music.
And it makes sense why. I feel like they will compete with even the no-name Dog, if the Dog launched a successful marketing video of an AI product.
I have seen this often in SF startups, heck I work for them, but man this is extreme.
But honestly all I see are long term price wars, I don't understand how this is a sustainable business strategy.
Or maybe that's the point... Who knows.
Havoc 8 hours ago [-]
Will stick to Jev. It’s cheaper, it’s the OG and we need diversity among provider
Imustaskforhelp 22 hours ago [-]
This rather didn't take long for OAI to create*, I remember people giving opinions and discussions that it won't take too long and that openAI should do it[0], so looks like they were right.
Interesting to see where all this leads us and if other major labs follow suit
Edit: decisions voice looks really interesting as well[1]
Decisions voice isn't a product for anyone else who was confused: it's a canned guide for hooking up a voice model to the decision model browser use thing
AM1010101 16 hours ago [-]
How well calibrated is it? Is 90% calibrated to be correct 9 times out of 10?
I think Jev had put significant effort here and its not clear if luna will be well calibrated in this way.
Conner_Hobbs 3 hours ago [-]
[flagged]
helloparallax 15 hours ago [-]
[flagged]
MiroslavPokorny 18 hours ago [-]
What value is there in knowing if a can has a dent ?
mikeryan 17 hours ago [-]
Manufacturing quality control. There are automated tools to pull things like dented cans off a line.
MiroslavPokorny 12 hours ago [-]
So AI is a trillion dollar industry and one of the biggest thing it can do is spot dents on a can ?
progx 10 hours ago [-]
Imagine a worker in a car factory who inspect the car's body.
Pictures from all angles and AI inspection is faster, cheaper more thorough and works 24 hours a day.
And that is only 1 scenario.
11 hours ago [-]
lab14 21 hours ago [-]
How is the pricing vs Jev?
jerrygenser 20 hours ago [-]
$0.10/mm input vs. $0.042/mm input. Both free output.
Topfi 20 hours ago [-]
In the same bench a full Jev run cost USD 0.0192,- vs Luna at USD 0.06,-, both via OpenRouter today. So about 3x in favour of Jev.
esafak 21 hours ago [-]
You knew it was going to happen! Benchmarks or it didn't happen.
krembo 14 hours ago [-]
But... Can it draw a pelican on bicycles?
Havoc 8 hours ago [-]
You probably could with the right harness. Like the logo turtle
OutOfHere 19 hours ago [-]
v3.26.0 of the openai Python SDK covers its use. Those already using the SDK don't need to make explicit HTTP calls.
waterTanuki 18 hours ago [-]
> The tulip became a luxury item and many varieties were introduced. The varieties were classified and the most sought-after, prized tulips were the streaked tulips, especially yellow or white streaks on a red or purple background. These flame-like tulips were highly sought after. Interestingly, the streaks or “flames” of the tulip petals were caused by a virus. The virus is the tulip breaking virus, or tulip mosaic virus.
No. The OpenAI subscription has never covered any API calls. The closest you can probably use via subscription is to get structured outputs via Codex.
hcentelles 18 hours ago [-]
[dead]
lin7c 15 hours ago [-]
[flagged]
chengjunsuger 14 hours ago [-]
[flagged]
steven_a 14 hours ago [-]
[dead]
MarrowAI 10 hours ago [-]
[flagged]
juanfranpaez 20 hours ago [-]
[flagged]
zane_shu 21 hours ago [-]
[flagged]
nnoman7808 15 hours ago [-]
[flagged]
dvt 20 hours ago [-]
I genuinely do not understand why anyone would pay OpenAI for this. Running something comparable to Jev is pretty trivial. The whole point of paying for ChatGPT is because OpenAI has a bunch of warehouses that can run a zillion-parameter model.
Running a decision model is way easier and much cheaper. Are they really just trying to capitalize on the hype here? It feels like they really have absolutely zero moat.
mediaman 20 hours ago [-]
Why would I run it myself? It's $0.10 per million tokens. Dirt cheap. (Jev is even cheaper.)
You could ask the same question about why anyone would rent a VPS. I can just run my own hardware, it's just a computer!
Buy vs rent is not just about what's possible, it's about what's economic.
lelandfe 18 hours ago [-]
For one, to remove network RTT
dvt 20 hours ago [-]
Yeah and OAI is twice as expensive as Jev, which is kind of my point. And more expensive than open models, which you don't necessarily have to host yourself. Pure bandwaggoning.
19 hours ago [-]
tmhall 20 hours ago [-]
For my use case it will cost like $11 a month and we already have OpenaAI keys and accounts with billing in place. I don't want to run my own model infra and I don't want to get permission to set up an account with typesafe.ai
csharpminor 20 hours ago [-]
If you're in an enterprise that already has a procurement agreement with OpenAI, this means you don't have to onboard another vendor. Bucket platform strategy.
drdexebtjl 19 hours ago [-]
My opinion is similar, but for a different reason: every use case for decision models that I can think of, I don’t want the model to change in X weeks when the lab decides to “improve it” or “make it safer”.
simonw 20 hours ago [-]
Depends on the quality of the results. These things are driven by text prompts. If it turns out the OpenAI one returns better quality results than open weight variants they'll be rewarded by the market.
Anyone using a decision model like this is going to have to spin up their own evals - these are far harder to vibe-check than regular text output LLMs.
TSiege 20 hours ago [-]
There isn’t a moat in the sense of self hosting but you need a reason for people who don’t want that to stay on your platform. Customers save time and effort managing payments easier this way. However it’s a race to the bottom price wise.
Going to be all about branding and platform stickiness for OpenAI to make investors and creditors whole.
super256 20 hours ago [-]
Existing enterprise contracts? Data retention contracts (some have zero data retention contracts)? Staying with a single provider because it's easier to have everything in one place?
There are probably a lot more reasons.
jcims 20 hours ago [-]
If you work for a company that has a 3 to 6 month onboarding period for new vendors and a lifetime commitment to maintain a whole bunch of vendor management horseshit for as long as that relationship exists, it makes a ton of sense.
Add in a bunch of model governance and oversight for anything you train yourself and it’s pretty much a slam dunk deal.
(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)
I know it's just a little typo but it made my morning :)
dejure: legal
defecto: enshittified
(adj.): The defecto way to playback music is Spotify.
I'm skeptical of the quality of probability calibration for models if you aren't giving them training data. The issue is that there is a prior distribution that's unique to your specific data and the calibration is really sensitive to that.
Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.
There are classes of problem where this shifts the economics from "paying a human to do this is cheaper than AI" to "the software is now cheaper than the human"
For instance if you have a predictive model for market prices that is not calibrated that's... nice. If you have a calibrated model you can add a Kelly better and you have a trading strategy that makes money. Similarly if you are classifying articles or images or other contents to make a feed you might believe that 70% or 95% or some other level of precision is "good enough" and you can set the knob and turn on the cruise control.
I've asked twice now about what I'm missing and for a specific use case where you can't just do this with a regular LLM call and nobody has replied that so if you have the answer that would be great. Looking for something specific instead of just it's faster or cheaper which is definitely nice but I'm just not seeing what this opens up that was not previously possible
And theoretically will give you better answers statistically as it's calibrated.
https://laya.convaiinnovations.com/
You can throw a user bio at it, like "NAME: John Smith, AGE: 71, LOCATION: California" and ask OpenAI:
* Is this user located in the United States?
* Is this user located on the East Coast?
* Is this user located on the West Coast?
* Is this user old enough to vote?
* Is this user old enough to retire?
4 out of 5 of those are all going to resolve Yes, with a far greater than 0.5 rate.
Toss a bunch of freeform text bios at it, and get folks categorized within any number of data-points you're looking for.
Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.
If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
and they ask dumb followup questions after 7 business days when you want different access
If you are charged based on usage, you can “soft switch” between them really quickly.
They even negotiate with all three aggressively.
So not really.
The AI companies want to differentiate and become something more than a commodity, even if it's as critical as a utility is.
Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"
Jev and this decisions api are mostly useful for inference at scale in a workload where cost and latency matter… and that’s where evals become crucial. Could coding tools use it? Sure, but that’s probably a special case.
This same suite can simply be run very handoff to switch to a new prod model.
If there's efforts to build vendor lock in, it's going to be in the surrounding api, whether it's streaming responses, or this weird prediction thing, or temperature settings, seeds, stuff like that.
And even then, competitors can copy the interface because interfaces are not copyrightable.
That shorthand rule of thumb is often applicable, but can lead to risky / incorrect assumptions.
The central distinction isn't simply interface versus implementation. It's functional systems and constraints versus protectable expression - a distinction that can cut through both an interface and its implementation.
And like Web 2.0 I’m sure the day will come where the walled gardens return. All the old SaaS companies are starting to charge for agentic access. Proportionally more of the AI budget will go to them instead of the model providers, who are still stuck in this price war
It's always been obvious that small models with a specific purpose will outplay the larger more general models. It becomes a question of "do I want to be able to toss _anything_ at this frontier model, have it handle it all but pay the price" or "do I want to toss _specific things_ at this micro model, have it handle it but not be flexible".
It's the same with agentic RAG/search, things are moving towards much smaller search-specific models rather than tossing RAG chunks at a large frontier LLM. It's like how I can ask a multimodal frontier model to identify bounding boxes for objects in an image...but if I want to do that faster/cheaper/at 60fps then I should be using yolo or similar.
Even for technical problems sometimes having something fast and convenient that doesn't require a ton of fine tuning opens up a lot of doors. Those door might've easily been openable previously with a few days of effort. But the difference between a few days and something that you can set up with a quick account and a prompt is massive
Hopefully people will flock to whatever product is making its mission to be commodity and the easiest to replace. Really don't want another free ingress, 100$/TB egress Cloud situation.
While the internet will live on, many of the companies that were initially leaders in it will not.
I’d bet by 2039 either Anthropic or OpenAI will have been bought by SpaceX. One will be the Sun Microsystems of this era.
Maybe they could for example point their amazing models to their GitHub issues, and use their power to actually fix the bugs and myriad of papercuts before solving cancer and world war/peace/whathever the stockholders believe they want.
> is .. all people want.
Try making more measured statements, you exaggerate your point and end up being wrong. Of course this new barely used product feature is not all people want. Of course this minor product feature is not the nail in the coffin of whatever argument, they are just responding fast to a smaller competitor by providing that feature themselves, something that happens thousands of times in business, did Instagram prove it was a commodity when it implemented reels by copying tik tok? Did Uber prove it was a commodity when it started offering food delivery? Doesn't make much sense.
That's what I have been saying: OpenAI or Anthropic's moat will not be the models, but how they integrated into enterprise clients and make it super hard for them to switch to a different provider. Kind of like Microsoft Office does it. Or Slack. Or Google Work. None of these products are better in any sense, they just cater to the needs of big enterprise clients in another way (customer relationship, sales, maybe compliance) and make it hard to switch
All of those companies that naive HNers used to say they could code in a weekend literally can be coded in a weekend now, so why haven’t they?
If enterprise is the goal they’re really at the whims of the old school software companies because they own no data of their own. Imagine Google decides to push Gemini one day and now ChatGPT can’t write docs anymore.
My approach would be per-user (or per-project) "memory".
I don't mean a tack-on like a RAG / document store with "notes" added to it in the background, but an actual medium- and long-term memory. Something like an extension of the KV cache stored in High Bandwidth Flash (HBF), or a subset of the weights trained "online", similar to LoRA.
This would not be transferable to any other base model, so would be excellent "lock in".
The downside is that the memories likely wouldn't be transferable to new models either, but I can imagine solutions to that too. I.e.: Train an MLP to "translate" from the old memory weight space to the new one.
You raise $200B to be a high margin low capex business, not an industrial commodity producer with high capex and margins being set by competitors who can duplicate your product and undercut you on price.
Clearly this was a rush job to respond to the competition. I am more curious about how the dedicated model will perform after they've had time to do it the right way. The probabilities I am seeing so far do not correspond with figures the business would find very agreeable.
The hidden danger with this could be demonstrating how thin the veil actually is. We may wind up reducing confidence in decisions simply by making their probabilities visible. Some kinds of information are quite hazardous.
* Using predicate questions gave nearly perfect/expected probability outcomes
* Asking it to choose an outcome behaved differently from drawing randomly - if the true probability of a red marble draw was 50%, using Decisions API produced 86%, i.e. it picked the right marble but gave a significantly more biased weight on its choice
* Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%
These models seem to have the same biases and limitations of LLMs minus speed. Outputs and inputs should be treated the same way as LLM prompts.
Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).
Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).
Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.
Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).
Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.
That's a big motivator.
How these models play out is an open question but existing provider contracts and T&C are important for enterprise.
Still surprised they even leveraged Luna for this. Given their resources in data, compute and manpower, would training a decision model from scratch take that much longer to not make sense given the cost, compute and performance advantages that would likely provide?
3 times more expensive at twice the latency with lower performance is a tough sell, though yeah, prior relationships will likely smooth some of those deficiencies over.
Mercury Decide approved some commands that weren't safe. It and Solar Decide were vulnerable to
Only Liquid D1 and Clef matched Jev's performance.For what it's worth, ran every task twice on each model, most were for some UI component synthesis and charting insanity that is a bit hard to explain, but some were simple tag selection, basic noul at threshold 60%. Essentially, whether to use the provided tag given the title of a browser tile:
Of course, tags can be a bit subjective, but in these cases, I'd argue the values provided by Jev were far more representative of my subjective assessment over Lunas. If SnP stuff on Bloomberg isn't investing, nothing is.Goal for tagging is mainly a near instant, over writable, sane default provided to users in the background. Resolve the whole "I love using Notion/Obsidian/PKM software of your choice but spend 80% of my time just thinking about the ideal tag before starting to read" issue. Lunas output is not really helpful here.
Also reran Jev [1] and Mercuy Decide [2] (which got 0.83) with same input, for reference.
Also, also, used the example via OpenRouter exactly as you did with the only change being my original input (which has my slightly odd tagging system and multiple tags in input though requests a noul as output) and got a 0.47 [3] on investing both times.
If Jev did well but both Mercury Decide and Luna failed, I'd chuck that up to my use/prompting, but Mercury Decide does well here so it seems Luna specific. Will add that I did also try Clef, happy to share that data if anyone wants, but quickly dropped Clef due to pricing vs Jev.
[0] https://imgur.com/a/nAkCZiG
[1] https://imgur.com/a/scIa9nF
[2] https://imgur.com/a/M7Mm1l7
[3] https://gist.github.com/Topfi/d77e503c7d1f6d11fc32d1b2174ec0...
In text mode I'd use structured output for this, or even get away with instructing the model to [emote] or even just map emojis to emotes.
Great demo though, I'm going to try out the image stuff now. Super rad that it's multimodal decision making; "does this pipe need to be inspected?"->"yes","I want a closer look","no" etc!
Edit: impressive stuff! I gave it a bunch of examples: - things that humans should/should not eat and it got this correct (including rejecting rat poison) - probability staff member should be called to a train platform (people standing vs. child looking over the edge)
It fails this example: "Given the context, and focusing on accuracy and efficiency, does the reasoning make sense."
Input: "Context: supplementary: {toolName: "get_weather", fetched: "2 minutes ago", currentWeather: "sunny and 25 degrees celcius"} user: hello how are you? assistant: I'm good, what can I do for you? user: what's the weather? Reasoning: the user is asking about the weather, so I should call the get_weather tool"
Output: 100% (expecting 0% since the requested data is available in context and not stale).
Prompt "Should the tool be called" also fails.
But the prompt: "Does the reasoning make sense given: - Facts and information available in the context - Requirements for efficiency in tool calling - Requirements for tool arguments to be provided only by the user" works pretty fine, it's mostly points 1 and 3 that give it the correct behaviour.
Fun stuff. I suppose for these decision models reasoning isn't enabled? It would explain why logic puzzles that require several steps to make a decision don't really do so well. It gets the carwash problem fine, but fails on logic problems such as selecting which word from a list where at least one letter appears in another word from the list.
https://arxiv.org/abs/2510.01238
Open source classifier models you can run and train locally on CPU
- Cost : It is the same for both scenarios $0.10 per 1M tokens
- Speed : decisions is 10x faster than responses API
- Quality : I guess if we compare with luna which is a pretty good model it itself, both will be at par
So essentially it has to do more with speed vs any other factor.
You can give it a CCTV image (like of a train platform) and ask it to quickly decide actions such as triggering an automated auditory alert, deferring to a larger model for more detailed analysis, etc.
Unfortunately, a deterministic prediction utterly fails to yield an uncertainty measurement which is critical to have in actual risk reduction. If you want the variance in measurement, it is vital to obtain multiple measurements. This also gives a confidence interval.
Luna is so cheap it's borderline free (without tool use), so I'm struggling to figure out where to use this/Jev.
For example product categorization. Why 'risk' using this/Jev when a Luna LLM call will be smarter (in theory)?
e.g. why return
and not (p' = 0.93 * p + 0.07 * 1/4)?
maybe to say something like "under 0.6 confidence, reject the answer anyway"
yea I think that's the benefit here
[1] https://github.com/tensorninja/ninfer-4090
I tried smaller decision models like laya (also with custom finetunes) but the accuracy was not really good (for the things I tested). Also I don't have any VRAM left on this GPU so i had to decide whether to host laya or qwen-3.8-27b but not both at the same time. Running decision models on a CPU will also be noticeably slower so I went down this route to have both combined with shared base weights.
[0] - https://news.ycombinator.com/item?id=49976996
but the decisions are stateless... or you plan on passing the last N frames?
Or does OpenAI start distilling their own models for use cases like decisions?
Also if you have a long “system prompt” then caching would have saved a considerable amount on bulk data processing.
There may well be a technical reason I don’t understand.
OpenAI thinks their API is really worth more than 2x the price?
Did you get that reversed? OpenAI has a ZDR guarantee while Anthropic doesn't.
https://docs.moondream.ai/
"If update from select group of individuals on WhatsApp is classified as urgent then interrupt my music."
Doable?
So definitely doable.
but about 10x faster inference
The fact that this keeps happening demonstrates there is no moat. The fact that each wave gets a little less attention demonstrates there is no killer product here.
Will this get folded into models / post training pipelines at some point and make them better at calibrated outputs?
I'm mostly waiting for EU endpoints
It seems this one caught openai on the back foot, and this is a scramble to maintain parity.
The company is clearly still innovating towards AGI rather than asking "what do people actually need?"
Despite once being the darling of AI it's
- lost it's models' performance edge, and got too many similar offerings
- continues to launch products without a market or isn't done better using other tools (e.g. dots)
- despite having AI can't lock down it's own products showing lack of skill
- focuses on solving maths problems humans can do for tests, when real world problems - disease, materials, energy research etc is all outstanding
- abandoned it's open model and open source programmes, despite Google, and multiple successful Chinese, and now European companies make their frontier models open weight.
- pissed off a portion of its non corporate fan base by killing GPT 4o instead of recognising the brand and product attachment as an opportunity
- launched laughing stock projects like being able to actually call chatgpt on a telephone number (wtf!)
- fails to capitalise on market segments like an AI that can provide corporate network sentry duties
It's increasingly looking like the company has jumped the shark and if I was an investor would be asking questions as to why it actually took so long to bring a jev like product to market, and why they are labelling something that is a simplification of existing models as "beta".
The whole point of AI as I see it is to make our life easier and answer the questions we can't. It isn't to make an AGI so powerful that it can replace us.
Along the way, that goal was forgotten, but it's not been forgotten by the new startups.
I am sure people will be able to use it, but if LLM progress is anything like before we will have Jev 2 in about a month.
Training a local model like Jev with some learnings that can be extremely cheap to host shouldn't take that long either.
So I honestly don't see the point of this, other than to put something out.
Although that does seem like OpenAI's strength turn around slop products and iteratively improve and try to out compete others in everything.
Only 2 things they have clearly given up on are Video models(no moat, copyright nightmare) and Music.
And it makes sense why. I feel like they will compete with even the no-name Dog, if the Dog launched a successful marketing video of an AI product.
I have seen this often in SF startups, heck I work for them, but man this is extreme.
But honestly all I see are long term price wars, I don't understand how this is a sustainable business strategy.
Or maybe that's the point... Who knows.
Interesting to see where all this leads us and if other major labs follow suit
Edit: decisions voice looks really interesting as well[1]
[0]: https://news.ycombinator.com/item?id=49802161: OpenAI is well positioned to fast-follow Jev
[1]: https://developers.openai.com/api/docs/guides/decisions-voic...
I think Jev had put significant effort here and its not clear if luna will be well calibrated in this way.
Source: https://www.canr.msu.edu/news/tulip_mania_the_history_of_the...
What's old is new.
Running a decision model is way easier and much cheaper. Are they really just trying to capitalize on the hype here? It feels like they really have absolutely zero moat.
You could ask the same question about why anyone would rent a VPS. I can just run my own hardware, it's just a computer!
Buy vs rent is not just about what's possible, it's about what's economic.
Anyone using a decision model like this is going to have to spin up their own evals - these are far harder to vibe-check than regular text output LLMs.
Going to be all about branding and platform stickiness for OpenAI to make investors and creditors whole.
There are probably a lot more reasons.
Add in a bunch of model governance and oversight for anything you train yourself and it’s pretty much a slam dunk deal.