Skip to content
Qila AI

02Security

Company documents should never travel through a personal AI account.

We run the GPUs and run the model on them. Every request passes your rules before it reaches that model, and the endpoint keeps neither the prompt nor the response.

01Zero retention

Zero-retention inference, stated plainly.

The model endpoint processes a request, returns the answer and retains neither the prompt nor the response. Prompts are not used for training.

Billing records token counts and user metadata, never prompt text. Workspace history is kept only if your company chooses, and each employee sees only their own.

What it does not mean

The model runs on GPUs that Qila runs and manages. The assistant application itself can be hosted inside your own infrastructure, so documents and conversation history stay on systems you control. We do not offer India data residency on the standard service. If your workload needs reserved capacity or a specific region, we scope a dedicated deployment separately.

Request trace
  1. ReceiveKept: Nothing

    Encrypted request from a signed-in employee

  2. CheckKept: Usage metadata

    User, department budget, approved model

  3. ProcessKept: Nothing

    Model runs on Qila-managed GPU compute

  4. ReturnKept: Token counts

    Answer goes back to the same employee

After the response

Prompt
Response
Billing
02What we promise

What we can promise. What we cannot.

Each row is backed in writing before a pilot starts.

GuaranteeStandard serviceDedicated deployment
Zero-retention inferencePrompts and responses are not retained by the model endpoint.
No training on your promptsSubmitted content is not used to train the models we run.
Encrypted in transitEvery hop between employee, Qila and model is encrypted.
Authenticated access onlyEach request comes from a signed-in employee.
No infrastructure credentials for staffKeys stay with Qila. Staff see only the workspace.
Company-level users, roles and budgetsSet by company, department or employee.
Minimal loggingToken counts and user metadata, not prompt text.
Reserved GPU capacityCapacity held for your company alone.
Chosen processing regionRegion fixed in the contract.
Model on GPUs Qila runsWe run the GPUs and run the model on them. Nothing goes to OpenAI or a consumer AI service.
Assistant hosted in your infrastructureThe workspace application runs on your servers. Documents and history stay with you; the model runs on GPUs we run.

Before a pilot starts

You receive four things in writing.

  1. 01The model endpoint's data-handling and retention terms, in writing.
  2. 02A list of exactly what Qila stores: accounts, usage metadata and, if you choose it, workspace history.
  3. 03The access model: who administers users, budgets and connected documents on your side.
  4. 04The history policy you want: none, per-employee, or reviewable by an administrator.

Next step

Test the controls on one sensitive workflow.

  1. 0130-minute call
  2. 02Fixed 2–4 week pilot
  3. 03Measured cost per task
  4. 04Monthly review