Skip to content

Products

Model endpoints

Use open AI models through an OpenAI-compatible API. Pay per token, keep every request inside Nepal, and never manage a GPU.

Early access is open to enterprises on NVIDIA H200. More GPUs, and access for more teams, are coming soon. We reply within one working day.

built for product teams shipping ai features. Technology: vllm · litellm · openai-compatible



Steady under real load

The serving engine (vLLM) batches requests continuously, so response times hold steady when many users arrive at once, not just in a benchmark.

Works with the OpenAI SDK

One OpenAI-compatible base URL for every model (a LiteLLM gateway). Point your existing SDK at it, change the model name, and you're done.

Pay per token

Input and output are priced separately, and repeated prompt prefixes cost less. Usage appears on the same rupee invoice as your GPU time.

Keys you control

Issue a key per service, limit it to specific models, rotate it without downtime and revoke it instantly. A key is shown once, when you create it.

Your prompts stay in Nepal

Prompts and responses are processed and logged inside Nepal. Nothing goes to a foreign model provider, because there isn't one.

Tools and images

Function calling across the catalogue, and image input on the models that support it.

np-ktm-1.corevalley.ai
$export CV=https://api.corevalley.ai
$curl $CV/v1/chat/completions \
-H "Authorization: Bearer $KEY" \
-d '{"model":"llama-3.3-70b",
"messages":[{"role":"user",
"content":"नमस्ते"}]}'
# served from np-ktm-1 · per token
$
API
OpenAI-compatible (LiteLLM)
Engines
vLLM · SGLang
Billing
per million tokens, in and out
Models
Llama · Qwen · DeepSeek · Mistral

Start with a quote.

Tell us what you want to run. We reply within one working day with a capacity plan and a firm rupee quote.


Other products

Keep exploring.