Private inference / workplace-native

Your AI. Your choice of cloud.

Bring AI into Slack, Teams, email, calendars, WhatsApp, Word, PowerPoint, IDEs, and CI. Launch a dedicated Ovrion Cloud workspace in minutes, or deploy the same gateway on infrastructure you control.

Choose a model, see compatible GPUs and the full hourly rate, then approve the first prepaid block. Stopping reconciles exact usage and returns unused reserved credit.

The request path

Trust is a route you can inspect.

A workplace request reaches your dedicated managed gateway or the gateway you operate, runs on the model endpoint you select, and returns with provenance. The deployment mode is explicit.

01 / SURFACE

Work starts where people already are.

Messages, mail, meetings, documents, source, and build logs.

02 / CONTROL

Your gateway authenticates and routes.

Signatures, OAuth, API keys, model aliases, billing gates, and provenance.

03 / EXECUTION

Your model endpoint does the work.

Local hardware, your private endpoint, or a GPU rented through Ovrion Cloud.

Safety, named precisely

Different safeguards deserve different words.

Ovrion does not turn code-level choices into OAuth-level promises. The product names where a boundary is externally enforced and where it is enforced by implementation.

Outlook mailDraft and summarizeMail.Send excluded from OAuth scopeENFORCED
GmailDraft and summarizeGoogle scopes also permit sendCODE-LEVEL
CalendarTentative RSVPNo organizer notification requestedENFORCED
WhatsAppOutbound replyCloud API has no draft responseEXPLICIT SEND

Outlook can structurally exclude mail sending at consent. Gmail cannot. Ovrion's Gmail client contains no send method, but Google grants broader capability to the authorized service account. WhatsApp replies require an outbound Meta API call by design.

Workplace surface

One boundary, many places to work.

Each surface maps to the same Ovrion gateway, managed or self-hosted, with the authentication and reply behavior appropriate to that platform.

Slack

Mentions, DMs, thread replies, slash commands, and interactive actions.

V0 HMAC / BOT LOOP GUARD Shipped

Microsoft Teams

Plain replies, Adaptive Cards, bot-added greetings, and mention stripping.

BOT FRAMEWORK JWT / AUDIENCE CHECK Shipped

Outlook

Draft replies, thread summaries, and silent tentative RSVP notes.

DELEGATED OAUTH / NO MAIL.SEND Shipped

Gmail + Calendar

Draft replies, summaries, and tentative RSVP notes through Workspace delegation.

ADMIN DELEGATION / CODE-LEVEL NO-SEND Release candidate

WhatsApp

Verified Business Cloud webhooks and real outbound text replies.

X-HUB-SIGNATURE-256 / META SEND Shipped

Word + PowerPoint

A live task pane that asks the gateway and inserts only when the user chooses.

OFFICE.JS / USER-CONTROLLED INSERT Live add-in

GPU economics

Know the rate before capacity starts.

Ovrion sizes a Hugging Face model with its context window, quotes the required GPU tier, and reserves the first billing block before launch. Every stop reconciles exact time to a named customer and GPU, then returns unused credit.

QUOTE / CONTEXT-AWAREBEFORE BOOT

Your selected model

Model weightsMeasured
KV cacheIncluded
GPU tierRight-sized
Billed rateVisible first
AUTO-STOP AT ZEROAPPEND-ONLY LEDGER

Commercial model

The engine is open. The workplace layer is licensed.

Run the core cluster freely. Subscribe when you connect workplace surfaces. Buy GPU credit only when customer hardware is not enough.

Core

$0

Open-source cluster, OpenAI-compatible gateway, local and private model endpoints.

Pro

$29 / month

Workplace integration access for one customer-run gateway.

Business

$99 / month

Business subscription record and priority-support tier. Final entitlement packaging remains subject to the published license terms.

GPU credit

Prepaid

GPU cost plus the Ovrion managed margin, reserved before each block and reconciled at stop. Campaigns may discount our margin; the total can never fall below the underlying GPU cost.

Managed or self-hosted

Start with the deployment that fits you.

Use Ovrion Cloud to choose a model and GPU without server setup, or pull the same published image onto hardware you control.