Jev was released on September 15. Within about 12 hours of testing it, our engineering team had the first Jev-powered workflow running inside Timeglass.
That speed came from knowing exactly where a model like Jev belonged.
Our CTO, Eddie, is probably the person who spends the most time testing models at Timeglass. So when Jev dropped, he heard about it almost immediately.
He was skeptical at first.
We had spent a lot of time tuning the models behind Timeglass. We knew what worked, what it cost, and where the tradeoffs were. Jev was a completely new kind of model, and we would have to rebuild our AI pipeline to use it.
But at Timeglass our users come first, so we still had to try it.
Jev is really good at decision making at scale
Most LLMs generate an answer one token at a time. Jev works differently. You give it context and a set of possible answers, and it returns probabilities for those options.
In Eddie's words, it's basically a very smart multiple-choice solver.
That happens to be extremely useful for Timeglass.
Timeglass takes your raw work activity and turns it into a structured record of your day. We watch your actions, group them into tasks, and then automatically turn them into timesheet entries, linking each of your tasks to a project, client, and category of work.
That last part involves a huge number of decisions.
Some of our customers have hundreds of active projects and clients, or dozens of categories and sub-categories. For everything each user does each day, we have to figure out where that time actually belongs.
This kind of decision making at scale is exactly what Jev is built to answer.
Jev beat our existing models
We have our own internal benchmarks at Timeglass, where we test every model that makes sense for our price, performance, and privacy requirements against real Timeglass workloads. This lets us stay on top of model releases and be confident that we’re delivering the best product for our users.
When Eddie ran Jev through those benchmarks, it made roughly 30% fewer serious classification errors than the models we were using before.
What counts as a serious error? For example, assigning work to the right client while failing to find the project is an error, but only a minor one. But assigning work to the wrong client or project can lead to fraudulent billing.
“We had to rebuild our timesheet prep pipeline for Jev, but it resulted in a 30% reduction in serious classification errors, while also being twice as fast.”
We also saw a meaningful improvement in processing speed. With Jev handling classification inside the pipeline, timesheet preparation now takes roughly 50% less time.
These improvements might sound incremental on their own, but they compound across an entire day of work. Every classification Timeglass gets right is one less thing someone has to review, correct, or think about when their timesheet is ready.
Jev went from release to production in days
Once the evals came back, Eddie rewrote the first pipeline around Jev. Roughly 12 hours later, we had the first implementation running internally.
We tested it on our own team, watched the output for a couple of days, and then shipped it to users.
Today, Jev helps Timeglass group individual actions into coherent blocks of work, assign tasks to projects, clients, and activity categories, and also resolve timesheet entry splits and merges.
This is pretty representative of how we build Timeglass.
There isn't one model underneath the product. Different models are good at different things, and our job is to keep figuring out which model should do which job. We continuously test models against our own workloads, with our own evals, under our own privacy requirements.
When something meaningfully improves the user experience, we ship it.
See how Timeglass turns a day of work into a timesheet ready for review.
Common Questions
FAQ
What is Jev?
Jev is an AI model released on September 15, 2026. Instead of generating an answer one token at a time, it takes context and a set of possible answers and returns a probability for each option.
What does Jev do inside Timeglass?
Jev groups individual actions into coherent blocks of work, assigns tasks to projects, clients, and activity categories, and resolves timesheet entry splits and merges.
How much did Jev improve Timeglass?
On internal benchmarks built from real Timeglass workloads, Jev made roughly 30% fewer serious classification errors than the previous models, and timesheet preparation now takes roughly 50% less time.
Does Timeglass rely on a single AI model?
No. Different models are good at different things. Timeglass continuously tests models against its own workloads, evals, and privacy requirements, and uses each one where it performs best.



