Ask
A small model on my desktop, answering for me
ConciergeConnecting…
Pick a topic — or just ask
Model—3 billion parameters, 4-bit quantised
Machinedesktop
GraphicsNoneevery token is computed on the CPU
Answer library—written ahead while the desktop is idle, and still growing
The trip your question takes
- 1Your browser → this siteVercel takes the question; it never touches the model.
- 2A queue in the middleThe desktop reaches out to collect it. Nothing can reach in.
- 3Safety railsPrompt-injection attempts, impersonation, and anything binding Bennett to work are answered by fixed rules — the model never sees them.
- 4The model reads a profileOne document about Bennett. If a fact is not in it, the honest answer is that it is not available.
- 5Back to youUsually 10–20 seconds. If it was asked before, it returns instantly.
The slow part is honest: a small model thinking on four CPU cores in a house, not a rented GPU. While nobody is asking, it works ahead — writing answers to likely questions so the common ones come back the moment you hit enter. Ask something new and you will feel the machine actually think.