Run aScribe on infrastructure you control. We install and configure it on your VPS or server, then you choose local engines for maximum privacy or cloud providers when you want speed without owning GPUs.
Configure transcription with WhisperX and summaries or library search with a local LLM via Ollama when your machine has the resources. Audio and text stay on your side. No cloud STT or cloud LLM bill for those jobs, and no third-party path for the raw content of those jobs.
WhisperX on CPU-capable hosts: workable without a GPU for many professional workloads
Ollama (or another local OpenAI-compatible endpoint) for summaries, chat, and embeddings when hardware allows
You pay for the VPS and power, not per-minute cloud transcription on that path
Built for legal, medical, and regulated teams that refuse silent data export
Ideal when the question is not only cost, but where the recording is allowed to travel.
How we configure your stack
aScribe is multi-provider by design. During installation we match engines to your hardware, latency needs, and data policy. Names below match what the product actually supports.
WhisperX runs locally (model sizes from tiny through large-v3-turbo). On a solid CPU host it is the practical on-prem path without a GPU: slower than cloud for long files, but your audio never leaves the box.
GPU available: local models
When the server has GPU capacity, we can enable heavier local models the product supports (for example WhisperX at larger sizes, plus NVIDIA Parakeet, Canary, or Voxtral depending on the build). Faster turnaround still stays on your hardware.
Cloud STT when it fits
If policy allows and you want throughput without GPUs, we wire AssemblyAI, Deepgram, or OpenAI Whisper with your keys. Useful for burst load; optional, not mandatory.
Summaries, chat, and library search
Local LLM: Ollama and friends
Summaries, per-recording chat, and embeddings can target a local OpenAI-compatible endpoint. Ollama is a common choice on a roomy VPS or workstation. Same machine, same trust boundary as the audio.
Cloud LLM when you prefer it
Need stronger models or less local RAM? Point the same features at a cloud OpenAI-compatible provider you choose. Embeddings can be configured independently of chat when that helps.
We do not force one vendor. The install conversation is about policy first: fully local (WhisperX + Ollama), hybrid, or cloud-assisted. Then we wire SMTP, storage, and handoff so your team can operate day to day.
Why self-host
Privacy, control, and long-lived archives on your host. Choose full local processing for sensitive work, or mix cloud providers when speed matters more than keeping every byte on-prem.
Licence note
Self-use under ELv2 is free. Re-offering aScribe as a hosted service to third parties requires a commercial licence from Majorum.
Installation on your server
From $799: we install, configure, and hand over a production-ready instance on infrastructure you control (for example your VPS).
Docker-based stack on your host as agreed
STT and LLM providers matched to CPU, GPU, or cloud keys you approve
Operational walkthrough for your team
Maintenance & support
Recurring help for upgrades, monitoring hooks, and questions. scoped and invoiced separately.
Tell us CPU vs GPU, expected languages and volume, and whether you need a fully local path (WhisperX + Ollama) or cloud assist. We will scope installation and timeline.