I built an on-device AI that judges every screen in a second

My thesis app needed the cloud; on-device, one screen took 20 to 40 seconds. A new kind of model cut that to one second, and I built Qualm for the Mac in a day.

Almost no screen-time tool uses AI. The idea is obvious: judge every screen instead of blocking a site. But that is a model call every few seconds on everything you read. In the cloud it costs money and privacy; on the device it was too slow.

I built the cloud version anyway. SeeNot, for Android, was, as far as I know, the first to judge the screen itself, and it became my thesis at SUSTech. It worked, and it never took off: people liked the idea, few downloaded it, and I never liked where the screenshots went.

On September 15, TypeSafe released Jev, a model that doesn't write text. You give it a text and a list of questions with fixed answers (yes or no, or one option from a list), and it answers all of them at once, in well under a second; because nothing is generated, ten questions take about as long as one. That is exactly what my app had been asking a VLM to do with a page-long prompt. And Kev, an open-source version of the same idea, does it on an Apple silicon Mac in about a second.

What was new was where it could run. The on-device model I had tried for SeeNot took 20 to 40 seconds per judgment; this takes about one, with nothing sent anywhere. So I built the second try, Qualm, a Mac menu bar app.

YouTube is a lecture and a Shorts feed; the lecture stays, and the Shorts get a pop-up that says why. Each new screen gets a short list of questions (what kind of page is this, what is it for, is it private, does it break each of my rules), and plain code decides. I no longer write a prompt; I choose the questions and write the code that acts on the answers.

Most of what the first version taught me is a list of what the tool must never do. It never hard-blocks. It runs local by default, and you can switch to TypeSafe's hosted Jev, but it never falls back to it quietly, because the tool reads everything you read. And rules are generalizable sentences, so "short videos made for endless swiping" catches a site nobody listed.

A week after the Jev launch, I sat down with it, ran trials the first day, had a working app that evening, and have used it daily since. I built it with Claude Code, and for agents too: every setting is also a command, so an agent can add a rule, test it on your own recent screens, and undo it. The night before shipping I asked for an audit as a fresh user and went to sleep; by morning 77 subagents had filed 204 problems, two of them serious.

You can try it at qualm.r-q.name, and the code is on GitHub.