Speak clearly in a quiet room for ~15 seconds — a few natural sentences.
Speak and it replies out loud. Or type in the box at the bottom. Two modes (side panel → Mode):
Your voice flows through five stages (shown in the bar at the top, and as faces in Voice mode):
Mic → Whisper → LLM → Voice → Speaker
Click + next to Voice, name it, and record 15s of natural sentences. It's cloned with F5-TTS (most accurate) or XTTS (faster).
Turn on Web search in the Tools panel to activate them. The model calls the right tool automatically — you can also just ask normally:
Turn on Settings → Desktop control and pumTALK can use your computer for you:
By default pumTALK only listens on your own machine. To reach it from anywhere, run it behind a reverse proxy (Nginx Proxy Manager, Caddy, etc.):
Upload PDF, Word, Excel, images and text files. They're indexed and the assistant uses them to answer your questions.
Off = STT & TTS run on CPU (NVIDIA stays free for the LLM). Turn on only if you have plenty of VRAM.
Restart pumTALK for a network-access change to take effect. Use "All interfaces" to reach it from other devices (pair with a token via PUMTALK_TOKEN for safety).