Products
Model endpoints
Use open AI models through an OpenAI-compatible API. Pay per token, keep every request inside Nepal, and never manage a GPU.
Early access is open to enterprises on NVIDIA H200. More GPUs, and access for more teams, are coming soon. We reply within one working day.
built for product teams shipping ai features. Technology: vllm · litellm · openai-compatible
Steady under real load
The serving engine (vLLM) batches requests continuously, so response times hold steady when many users arrive at once, not just in a benchmark.
Works with the OpenAI SDK
One OpenAI-compatible base URL for every model (a LiteLLM gateway). Point your existing SDK at it, change the model name, and you're done.
Pay per token
Input and output are priced separately, and repeated prompt prefixes cost less. Usage appears on the same rupee invoice as your GPU time.
Keys you control
Issue a key per service, limit it to specific models, rotate it without downtime and revoke it instantly. A key is shown once, when you create it.
Your prompts stay in Nepal
Prompts and responses are processed and logged inside Nepal. Nothing goes to a foreign model provider, because there isn't one.
Tools and images
Function calling across the catalogue, and image input on the models that support it.
- API
- OpenAI-compatible (LiteLLM)
- Engines
- vLLM · SGLang
- Billing
- per million tokens, in and out
- Models
- Llama · Qwen · DeepSeek · Mistral
Start with a quote.
Tell us what you want to run. We reply within one working day with a capacity plan and a firm rupee quote.
Other products