Skip to content
Qila AI

04Pricing

Price the work, not the headcount.

No per-seat charge. A service fee, compute at published rates, and a cap you set.

01How billing works

Three lines on the invoice. No hidden fourth.

Priced as a managed service, not as software seats. You set the cap; we optimise the cost on our side. Compute is itemised at published model rates so you see exactly what the infrastructure cost.

Sample invoiceSample month · 40 active users
01Managed serviceMonthly fee, scoped to rolloutScoped
02GLM-5.3-Flash compute2,310 M tokens at published rates, no markup₹12,400
03Cap set by youUsage paused above this amount₹15,000
No per-seat charge · no markup on computePilot billed separately as a fixed fee
PilotFixed fee
Two to four weeks, one workflow, one team, up to 25 employees. Scoped after a 30-minute discovery call.
ManagementMonthly fee
Provisioning, monitoring, model updates, onboarding, support and the monthly cost review. Scoped to your rollout.
ComputeAt cost
GLM model usage itemised at published token rates, with no markup, under a hard monthly cap you set.
02Compute planner

Tell us what the AI will actually do.

Different work consumes compute differently. Answer a few concrete questions and the planner returns a deliberately cautious monthly GLM budget, not a best-case teaser.

Default model · Managed GPU inference

GLM-5.3-Flash

Strong at coding, tool use and long documents, with a 1M-token context window. You pay for tokens processed, itemised at the published rates on the right.

Prompt

$0.45 / 1M tokens

Cached prompt

$0.09 / 1M tokens

Completion

$1.50 / 1M tokens

zai-org/GLM-5.3-FlashRates checked 1 Sep 2026 · converted at ₹88/USD

Question 1 of 2

What kind of work are you trying to do?

Choose everything that applies. Human conversations, coding loops and unattended agents are kept separate because one “heavy user” label hides too much.

Included in this number
GLM-5.3-Flash prompt, cached-prompt and completion tokens at published rates, converted at ₹88/USD.
Not included
Pilot or management fee, storage, search tools, document parsing, taxes or reserved GPU capacity.
03Ways to start

Three ways to work with Qila.

Most companies begin with the pilot, keep the managed service and move to dedicated infrastructure only when the workload calls for it.

Option

Secure AI Pilot

Test one workflow with one team before committing to a rollout.

  • Branded private workspace
  • One knowledge workflow connected
  • Up to 25 employees
  • Usage dashboard and weekly readout
  • Fixed fee, scoped after discovery
  • Model usage at cost
  • No rollout commitment
Plan my pilot

Core service

Managed Private AI

Keep the assistant running after the pilot and extend it to more departments.

  • Everything in the pilot
  • All-staff rollout and training
  • Departmental budgets and access control
  • Monthly usage and cost review
  • Priority support
  • Monthly fee, scoped to your rollout
  • Model usage itemised at cost
  • Caps set by department
Plan the rollout

Option

Dedicated Infrastructure

For workloads that need reserved capacity, a chosen region or isolated processing.

  • Reserved GPU capacity
  • Region defined up front
  • Custom SLAs
  • Deeper integrations with your systems
  • Quoted after technical review
  • Reserved compute at cost
  • Capacity and region in the contract
Review my requirement
04Questions

Asked by owners. Answered plainly.

Where does the model actually run?

We run the GPUs and run the model on them, on cloud infrastructure Qila manages. Nothing goes to OpenAI or a consumer AI service. The assistant application can also be hosted inside your own infrastructure so documents and history stay with you. Reserved capacity or a specific region is scoped as a dedicated deployment.

Does the model endpoint retain our prompts?

No. It returns the answer and retains neither the prompt nor the response, and prompts are not used for training. Your workspace may keep chat history if your company chooses to. Both policies are documented in writing before the pilot.

Can spending be capped?

Yes. Limits are set by company, department or employee. When a limit is reached the workspace pauses access until an administrator changes it.

Why is the compute estimate deliberately high?

A low estimate is useless if the company hits its cap mid-month. The planner assumes a busy month, conservative cache rates and a 35% reserve. After the pilot we replace those assumptions with your measured cost per task.

What does the compute estimate include?

GLM-5.3-Flash prompt, cached-prompt and completion tokens at published rates. It excludes the pilot and management fees, storage, search tools, document parsing, taxes and reserved GPU capacity.

How does this compare with ChatGPT Business or Microsoft Copilot?

Those sell per-seat software on infrastructure you do not control, and the bill grows with headcount. Qila runs the GPUs and the model, connects approved documents, gives you access and budget controls, and supports employees. You control the spend; we optimise the cost, so a firm of 25 or more usually pays less than it would for individual ChatGPT subscriptions.

How long does a pilot take?

Two to four weeks, depending on the document source, the access requirements and the workflow you choose.

Next step

Start with one workflow and one team.

  1. 0130-minute call
  2. 02Fixed 2–4 week pilot
  3. 03Measured cost per task
  4. 04Monthly review