Rendered at 12:36:17 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
4 days ago [-]
ricardobeat 4 days ago [-]
This will only work for simple text classification tasks, which is the least interesting possible use of Jev.
tgluck 4 days ago [-]
Obviously it only helps when the same question is asked many times, but that's the case it's built for, and the case I have, every Jev question in my other project repeats thousands of times, and most Jev uses I've seen online look the same
gingersnap 4 days ago [-]
Is the local model similar to model2vec?
tgluck 4 days ago [-]
Not really
they distill different things. model2vec distills a sentence transformer into static embeddings, so the output is a faster general-purpose encoder.
Jevstiller keeps the encoder frozen (bge-small by default) and distills Jev's decisions on one specific question into a small head on top of it
4 days ago [-]
tgluck 4 days ago [-]
Author here. This puts a proxy in front of repeated Jev classification calls. At first everything goes to Jev; from Jev's answers it trains a small head on frozen sentence embeddings, picks a confidence threshold with an exact finite-sample bound so that at most 2% of all requests get an answer Jev wouldn't have given, and then answers the confident share locally at ~15 ms on a CPU. A permanent 2% audit keeps checking; if agreement breaks, everything falls back to Jev and it retrains.
Known limits: agreement is not accuracy (if Jev is wrong, so is the local model); coverage tracks how consistent Jev itself is (22% on noisy tweet tasks, 80% on news); it speaks Jev's API only, an OpenAI-compatible front is on the roadmap. Since 0.4.0 the guarantee can also cover "would Jev have been unsure", which matters if your code routes low-confidence answers to review. Apache 2.0.
kodefreeze 4 days ago [-]
Isn't this against their ToS? Useful for hobby stuff.
KetoManx64 8 hours ago [-]
Why would it be against the rules to use the previous answers that the AI model gave you within your own project? That's like saying you can't use your Claude code convo history to answer questions within your codebase.
tgluck 4 days ago [-]
[dead]
dotancohen 4 days ago [-]
It would be great if we could correct Jev's incorrect answers, even on a separate endpoint. Let me tell it what Jev got wrong.
What type of head is that? What type of model is that head part of?
tgluck 4 days ago [-]
Not today, but Interesting idea. The main motivation was a drop-in for an existing Jev setup, so the only teacher right now is Jev and the audit measures agreement with Jev. A correction would have to become a second label source that overrides Jev's for that input.
The head is a multinomial logistic regression: one linear layer plus softmax on top of a frozen sentence-embedding model (bge-small by default, swappable). That head is the entire local model, the encoder is off the shelf and never changes.
dotancohen 4 days ago [-]
Yeah, I kinda figured that head was the whole model, the way you phrased it. scikit-learn?
tgluck 4 days ago [-]
No, plain numpy. It's full-batch Adam on cross-entropy against soft targets, about 80 lines.
dotancohen 3 days ago [-]
I'll look into Adam, thank you!
wedg_ 4 days ago [-]
Woah cool idea. So it's almost a drop-in replacement for a typical Jev setup that just reduces your jev bill over time ?
tgluck 4 days ago [-]
Thanks.
Drop-in yes: point TYPESAFE_BASE_URL at it and nothing else changes.
they distill different things. model2vec distills a sentence transformer into static embeddings, so the output is a faster general-purpose encoder.
Jevstiller keeps the encoder frozen (bge-small by default) and distills Jev's decisions on one specific question into a small head on top of it
Known limits: agreement is not accuracy (if Jev is wrong, so is the local model); coverage tracks how consistent Jev itself is (22% on noisy tweet tasks, 80% on news); it speaks Jev's API only, an OpenAI-compatible front is on the roadmap. Since 0.4.0 the guarantee can also cover "would Jev have been unsure", which matters if your code routes low-confidence answers to review. Apache 2.0.
What type of head is that? What type of model is that head part of?
The head is a multinomial logistic regression: one linear layer plus softmax on top of a frozen sentence-embedding model (bge-small by default, swappable). That head is the entire local model, the encoder is off the shelf and never changes.
Drop-in yes: point TYPESAFE_BASE_URL at it and nothing else changes.