Ask
A small model on my desktop, answering for me
ConciergeConnecting…
Pick a topic — or just ask
Model3 billion parameters, 4-bit quantised
Machinedesktop
GraphicsNoneevery token is computed on the CPU
Answer librarywritten ahead while the desktop is idle, and still growing
The trip your question takes
  1. 1Your browser → this siteVercel takes the question; it never touches the model.
  2. 2A queue in the middleThe desktop reaches out to collect it. Nothing can reach in.
  3. 3Safety railsPrompt-injection attempts, impersonation, and anything binding Bennett to work are answered by fixed rules — the model never sees them.
  4. 4The model reads a profileOne document about Bennett. If a fact is not in it, the honest answer is that it is not available.
  5. 5Back to youUsually 10–20 seconds. If it was asked before, it returns instantly.

The slow part is honest: a small model thinking on four CPU cores in a house, not a rented GPU. While nobody is asking, it works ahead — writing answers to likely questions so the common ones come back the moment you hit enter. Ask something new and you will feel the machine actually think.

ProjectsContact