Deploy your way
A single-tenant instance we operate in the region you choose, or the full stack on your own infrastructure with our support. Same governed memory runtime; the model is yours to choose.
Choose your deployment model
The standard offer is an instance we operate for you, kept current and supported. Self-hosting is available where policy requires it, as a premium arrangement with support and updates agreed separately. Both run the identical runtime: memory, governance gate, audit, and learning loop.
Yohanun Cloud
The standard offer · invite-only
A dedicated, single-tenant instance in the region you choose, EU or UK, run by us and kept current. Nothing of yours on a shared instance. We work closely with each team we onboard; access is currently by invitation while we grow deliberately.
Perfect for:
- Validating your use case fast
- Teams without infrastructure appetite
- Design partners who want direct access to us
What you get:
- Zero infrastructure management
- Your region: EU or UK, single-tenant
- Updates and improvements as they ship
- Usage dashboard and full REST/WebSocket API
Self-Hosted
Premium · full control
The entire platform on your infrastructure: your network, your disks, your keys, and if you want it, your own model. An instance behind your firewall is one we cannot patch or monitor unaided, so self-hosting comes with a support and update arrangement agreed separately. It is the premium option, and a first-class one.
Perfect for:
- Regulated industries with data residency requirements
- Privileged data that can't leave your perimeter
- Teams that want the governance layer under their own roof
What you get:
- Complete data sovereignty
- Docker Compose stack: one command up
- Bring your own model: a locally served Llama or Mistral through any OpenAI-compatible endpoint
- Deployed, updated and supported under a separate agreement, with a measured go-live for your model
The model is a separate choice
Chosen per deployment, per app, or per request, and changeable later without losing anything: the memory, the labels, the ledgers and the lessons all live below the model.
A provider's model
OpenAI, Anthropic, Google, DeepSeek or Mistral under enterprise terms that exclude training on your inputs. The default, and the fastest to start.
Your own keys
The same providers under your account, stored for your tenant only, so usage and terms sit with you. An option on any deployment.
Your own model
An open-weight model such as Llama or Mistral served on your infrastructure, so not even a query leaves your control. A premium arrangement, agreed separately, with a measured go-live; any model training is a separate matter.
What's under the hood
The same five-service stack in both deployment models
The Stack
Coordinated Stores
- • FastAPI runtime: REST, SSE streaming, WebSocket
- • PostgreSQL: structured data, rules, audit ledgers
- • Qdrant: vector memory, where the access gate runs
- • Neo4j: entity & relationship graph, conflict walls
- • Redis: cache, sessions, instant revocation
API Integration
curl https://app.yohanun.com/api/ai/chat \
-H "X-API-Key: $YOHANUN_KEY" \
-H "Content-Type: application/json" \
-d '{"message": "What did we decide last week?",
"user_id": "user_123"}'
Self-Hosted Requirements
System Requirements
- • 4+ CPU cores, 8GB+ RAM minimum
- • 50GB+ storage (SSD recommended)
- • Docker & Docker Compose
- • Outbound access to your model provider (or local models)
Deployment
# The whole platform, one command docker-compose up -d # Validated deploys with staging + rollback make validate-deployment make deploy-staging && make test-staging
Runs Anywhere Docker Does
No lock-in, in either direction
Start on cloud and move on-premises when compliance demands it, or start self-hosted from day one. Your memory, rules, and audit history are your data. And because the runtime is model-agnostic, you're never locked to a model vendor either.
Ready to deploy governed memory?
Tell us about your environment and constraints and we'll recommend the right deployment model and walk you through it.