Self-hosted

Run aScribe on infrastructure you control. We install and configure it on your VPS or server, then you choose local engines for maximum privacy or cloud providers when you want speed without owning GPUs.

Secure private server and waveform, self-hosted aScribe

Strongest privacy path

Near-zero variable cost, complete data control

Configure transcription with WhisperX and summaries or library search with a local LLM via Ollama when your machine has the resources. Audio and text stay on your side. No cloud STT or cloud LLM bill for those jobs, and no third-party path for the raw content of those jobs.

  • WhisperX on CPU-capable hosts: workable without a GPU for many professional workloads
  • Ollama (or another local OpenAI-compatible endpoint) for summaries, chat, and embeddings when hardware allows
  • You pay for the VPS and power, not per-minute cloud transcription on that path
  • Built for legal, medical, and regulated teams that refuse silent data export

Ideal when the question is not only cost, but where the recording is allowed to travel.

How we configure your stack

aScribe is multi-provider by design. During installation we match engines to your hardware, latency needs, and data policy. Names below match what the product actually supports.

Engines we configure with you

WhisperXGPU localAssemblyAIDeepgramOllamaOpenAI-compatible

Speech-to-text

CPU VPS: WhisperX

WhisperX runs locally (model sizes from tiny through large-v3-turbo). On a solid CPU host it is the practical on-prem path without a GPU: slower than cloud for long files, but your audio never leaves the box.

GPU available: local models

When the server has GPU capacity, we can enable heavier local models the product supports (for example WhisperX at larger sizes, plus NVIDIA Parakeet, Canary, or Voxtral depending on the build). Faster turnaround still stays on your hardware.

Cloud STT when it fits

If policy allows and you want throughput without GPUs, we wire AssemblyAI, Deepgram, or OpenAI Whisper with your keys. Useful for burst load; optional, not mandatory.

Summaries, chat, and library search

Local LLM: Ollama and friends

Summaries, per-recording chat, and embeddings can target a local OpenAI-compatible endpoint. Ollama is a common choice on a roomy VPS or workstation. Same machine, same trust boundary as the audio.

Cloud LLM when you prefer it

Need stronger models or less local RAM? Point the same features at a cloud OpenAI-compatible provider you choose. Embeddings can be configured independently of chat when that helps.

We do not force one vendor. The install conversation is about policy first: fully local (WhisperX + Ollama), hybrid, or cloud-assisted. Then we wire SMTP, storage, and handoff so your team can operate day to day.

Why self-host

Privacy, control, and long-lived archives on your host. Choose full local processing for sensitive work, or mix cloud providers when speed matters more than keeping every byte on-prem.

Licence note

Self-use under ELv2 is free. Re-offering aScribe as a hosted service to third parties requires a commercial licence from Majorum.

Installation on your server

From $799: we install, configure, and hand over a production-ready instance on infrastructure you control (for example your VPS).

  • Docker-based stack on your host as agreed
  • STT and LLM providers matched to CPU, GPU, or cloud keys you approve
  • Operational walkthrough for your team

Maintenance & support

Recurring help for upgrades, monitoring hooks, and questions. scoped and invoiced separately.

Start an installation conversation

Tell us CPU vs GPU, expected languages and volume, and whether you need a fully local path (WhisperX + Ollama) or cloud assist. We will scope installation and timeline.