What is private LLM deployment?
Private LLM deployment is the practice of running an open-weight large language model on infrastructure your organisation controls, whether a cloud account, a rented GPU server or your own hardware, so that the text you send it is processed without passing through an external AI provider. The model weights are downloaded once; every request after that stays inside your network.
The pieces are straightforward to name. An open model such as a Llama, Mistral or Qwen variant. One or more GPUs with enough memory to hold it. An inference server that loads the model and answers requests efficiently. A private network so nothing outside can reach it. And the applications that call it: an internal chat tool, a document assistant, a classifier inside your workflow.
What makes it hard is not any single piece but the combination: choosing a model your GPUs can actually run at acceptable speed, keeping the whole chain private including logs and backups, and knowing honestly how the answers compare with what a hosted API would give you.
Organisations usually arrive here for one of three reasons: contracts or regulators that restrict sending data to third parties, a very high volume of requests where per-token pricing becomes expensive, or a desire to avoid depending on a single AI vendor's pricing and policy changes.
- An open-weight model with a licence that fits your use
- GPU capacity sized to the model and your traffic
- An inference server with batching and an internal API
- A private network, access control and careful logging
- Applications that call the model for real work
Do you really need a private LLM, or will a hosted API do?
You need a private LLM when a contract, regulator or internal policy forbids sending the data to a third-party AI service, when your request volume makes per-token pricing clearly more expensive than renting GPUs, or when you must run in an isolated network. If none of those applies, a hosted API is often the better choice.
Hosted APIs have moved a long way on privacy. OpenAI's API documentation, for example, states that since 1 March 2023 data sent to its API is not used to train its models unless you opt in, that abuse-monitoring logs are kept for up to 30 days by default, and lists India among the regions available for data residency. For many businesses, that combination satisfies their policies without the effort of self-hosting.
Be precise about what “data never leaves our control” means for you. A private model protects prompts and answers from an AI vendor. It does not protect them from a leaky application log, a monitoring tool that ships traces to an overseas service, a backup bucket with public access, or staff who copy answers elsewhere. We map every place data travels before recommending anything.
The honest answer is sometimes a hybrid: sensitive workflows on a private model, everything else on a hosted API. We design the internal API so applications can switch between the two without code changes.
Private deployment makes sense when
Contracts or policy forbid third-party AI processing, volume is high and steady, or the network must be isolated.
A hosted API makes sense when
Data rules allow it, volume is modest or bursty, and you want the strongest models without running GPUs.
A hybrid makes sense when
Only some workflows touch restricted data, and the rest benefit from top hosted models.
Choosing a model for private LLM deployment: Llama, Mistral, Qwen or an Indic model
Pick the smallest open model that meets your quality bar on your own test set, under a licence your lawyers accept, with good handling of the languages your users write in. For many business tasks, a model in the 7–14 billion parameter range is enough; larger models help with complex reasoning but cost far more to serve.
Licences differ. Mistral-7B-Instruct-v0.3 and Qwen2.5-7B-Instruct are released under Apache 2.0, according to their model cards. Meta's Llama 3.1 Community Licence allows commercial use but requires organisations with more than 700 million monthly active users to request a licence from Meta, and asks users to display “Built with Llama”. Your counsel should read the licence for the exact model you deploy; we do not give legal advice.
Context length and safety features matter too. Qwen2.5-7B-Instruct's model card lists support for up to 131,072 tokens of context, useful for long documents. Mistral's model card for its 7B instruct model notes that it has no built-in moderation mechanisms, which is a reminder that guardrails in a private deployment are your job.
For Hindi and Hinglish, test Indic-focused models alongside general ones. Sarvam-M's model card describes a 24-billion-parameter model built on Mistral Small, under Apache 2.0, with support for Indic scripts and romanised Indian languages. Our Hindi AI chatbot page covers how we test language quality.
How much GPU memory does a private LLM need?
As a rough rule, a model's weights need about two bytes per parameter at 16-bit precision, so a 7–8 billion parameter model needs around 15–16 GB for weights alone, plus headroom for the cache that grows with context length and the number of simultaneous users. Four-bit quantisation cuts the weight memory to roughly a quarter.
That headroom is where many plans go wrong. Each active conversation keeps a key-value cache in GPU memory, and its size grows with the length of the prompt and answer. A model that fits comfortably when one person asks short questions can run out of memory when twenty staff paste long documents at once. We size for your expected concurrency and context, not for a single demo request.
Practical pairings look like this. A 7–8B model at 16-bit fits on a single 24 GB GPU for modest context and a handful of users. A 13–14B model at 16-bit needs about 28 GB for weights, so either two 24 GB GPUs or one larger card. A 24B model at 16-bit needs around 48 GB. A 70B model at 16-bit needs roughly 140 GB, meaning several data-centre GPUs, or about a quarter of that with 4-bit quantisation.
These are planning estimates, not guarantees. We confirm them with a load test on the actual instance before you commit to a long rental.
- Weights: about 2 bytes per parameter at 16-bit, about 0.5 at 4-bit
- Add headroom for the key-value cache per active request
- More users and longer context need more memory
- Always load-test on the real instance before committing
Quantisation: running a bigger model on smaller GPUs
Quantisation stores model weights at lower precision, such as 8-bit or 4-bit instead of 16-bit, so the model needs less GPU memory and often runs faster, at the cost of some quality. For private LLM deployment on a budget, it is frequently the difference between one GPU and three.
The quality cost varies by model, method and task. Eight-bit formats are usually close to the original. Four-bit formats such as AWQ or GPTQ lose more, and the loss tends to show on harder reasoning, maths and less common languages before it shows on simple drafting or classification. A 4-bit larger model can still beat a 16-bit smaller one, which is why testing both is worth an afternoon.
Inference servers support many formats. vLLM's documentation lists FP8, INT8, INT4, GPTQ, AWQ and GGUF among the quantisation methods it handles, so the choice is rarely blocked by tooling.
We run your test set on each candidate: the full-precision model, an 8-bit version and a 4-bit version, and show accuracy, speed and memory side by side. You pick the trade-off with the numbers in front of you. For Hindi or Hinglish workloads, we pay particular attention to the 4-bit results, since lower-resource languages can degrade first.
Private LLM deployment on AWS in India: Mumbai and Hyderabad
AWS has two regions in India: Asia Pacific (Mumbai), ap-south-1, which is enabled by default, and Asia Pacific (Hyderabad), ap-south-2, which AWS lists as an opt-in region you must enable first. Deploying in either keeps the model, data and logs on infrastructure located in India.
For GPU instances, AWS describes its G5 family as using NVIDIA A10G Tensor Core GPUs with 24 GB of memory each, from one GPU on g5.xlarge up to eight GPUs on g5.48xlarge. That range covers most small and mid-sized private models. Larger models need instance families with bigger data-centre GPUs; availability of specific GPU types varies by region and changes over time, so we check live availability and quotas in your account before planning.
Within the region, the model runs in a private subnet with no public IP. Applications reach it through an internal load balancer or private endpoint, and access is limited to specific services and roles. Model weights sit in encrypted storage in the same region, and logs go to a log service in the same account with retention you set.
GPU quotas are a practical hurdle: new AWS accounts often need to request higher limits for GPU instances, which can take time. We start that request on day one. Cost controls matter too, because an idle GPU instance bills by the hour; scheduling off-hours shutdowns for internal tools can save a large share of the bill.
Indian GPU clouds and your own servers
Indian GPU cloud providers and on-premise servers are the two alternatives to the big global clouds for private LLM deployment. Indian providers can offer larger data-centre GPUs with data held in India; on-premise gives the most control but moves hardware, power and cooling onto you.
One example is E2E Networks, an NSE-listed Indian cloud provider whose website lists NVIDIA H100, H200 and A100 80 GB GPUs among its offerings. Providers like this suit organisations that want high-memory GPUs, billing in India and data kept in the country. We compare them with AWS on GPU availability, network setup, support and total monthly cost for your specific model and traffic.
On-premise servers make sense when policy requires a fully isolated network or when steady, high utilisation over years makes buying cheaper than renting. We set up the software stack on servers you provide or buy: drivers, containers, the inference server, networking and monitoring. We do not supply, install or maintain physical hardware, and we do not visit sites; your IT team or hardware vendor handles racking and power.
Whichever you choose, we keep the stack portable: the same container images and configuration run on AWS, an Indian GPU cloud or your own machines, so moving later is a migration, not a rebuild.
Inference servers: why vLLM is our usual starting point
An inference server loads the model onto the GPU and answers requests efficiently, handling many users at once. vLLM is our usual starting point for private LLM deployment because it combines high throughput with an API that existing applications already understand.
vLLM's documentation describes it as a library for LLM inference and serving that originated at UC Berkeley's Sky Computing Lab. Its key techniques are PagedAttention, which manages the attention cache memory efficiently, and continuous batching, which slots new requests in as others finish rather than waiting for a fixed batch. It offers an OpenAI-compatible API server and supports over 200 model architectures from Hugging Face.
The OpenAI-compatible API is more useful than it sounds. Applications written for a hosted API can point at your private endpoint by changing a base URL and model name. That makes hybrid setups and later migrations easy, and it means your developers do not have to learn a new interface.
For very small models, laptops or quick experiments, lighter tools such as llama.cpp or Ollama are convenient. For production with multiple users, we prefer a server designed for concurrency. Either way, we put an authentication layer in front, because inference servers are built for speed, not for access control.
Private LLM deployment and the DPDP Act: data residency in India
The Digital Personal Data Protection Act, 2023 governs how organisations in India process digital personal data, and a private LLM deployment can support your obligations by keeping processing within infrastructure you control, limiting access and recording what happens. It does not make you compliant on its own; compliance depends on your purposes, consent, notices and processes.
The Act received presidential assent on 11 August 2023, and its provisions are being brought into force in phases, with the Data Protection Board and core provisions active from 13 November 2025 and remaining provisions following through 2026 and 2027. On transfers outside India, the Act takes a negative-list approach: transfers are generally allowed except to countries the Government restricts. Sector rules from regulators can be stricter than the Act, so check those too.
What we build to support you: hosting in an Indian region, network isolation, encryption at rest and in transit, role-based access, masking of personal identifiers in logs, retention settings you define and audit logs of who called the model. We also map every data flow, including monitoring and backup tools, so there are no surprise transfers.
We are developers, not lawyers. Consent wording, notices, retention periods and whether a given flow is permitted are decisions for you and your legal adviser, and we are happy to work alongside them.
- Hosting in AWS Mumbai or Hyderabad, or an Indian provider
- No public endpoint; private network only
- Encryption at rest and in transit
- Masked identifiers in logs, with defined retention
- Audit trail of which service or user called the model
Securing a private LLM: network, access, logs and prompt injection
Secure a private LLM the way you would a database holding sensitive records: no public access, strong authentication, minimal permissions, encrypted storage and careful logging. The model itself adds one new risk, prompt injection, which needs its own defences.
Network first. The inference server sits in a private subnet with no inbound internet access. Only named services can reach it, through an internal endpoint. Outbound access is blocked or restricted, so a compromised component cannot quietly send data out. Administrative access goes through a bastion or session manager, never open SSH.
Logs are the most common leak. Default configurations in many tools log full prompts and responses. We decide deliberately what to log: request metadata, token counts and latency always; full text only where you need it for quality review, with personal identifiers masked and retention limited.
Prompt injection happens when text the model reads, such as an uploaded document or an email, contains instructions that try to override yours. A private model is just as vulnerable as a hosted one. We limit what the model can do (no direct write access to systems without confirmation), keep system instructions separate from user content and test with hostile inputs before launch.
- Private subnet, no public IP, restricted outbound
- Authentication in front of the inference server
- Least-privilege roles for every calling service
- Deliberate logging with masking and retention
- Prompt-injection tests on documents and emails
Private LLM vs hosted API: how much quality do you give up?
You usually give up some quality on hard, open-ended reasoning, and often very little on narrow, well-defined tasks such as classification, extraction, summarising a known document type or drafting from a template. The only honest way to know for your case is to measure both on your own test set.
The largest hosted models are trained and served at a scale no single business can match, and they tend to lead on complex analysis, long multi-step instructions and less common languages. Open models of the size most businesses can afford to run have closed much of the gap on everyday tasks, and a model that is slightly weaker in general can be equal or better on a specific task once it has good retrieval and, where justified, a fine-tuned adapter.
We build a test set of 50 to 200 real tasks from your workflows, with expected answers or grading rules, and run it on a hosted API baseline and on two or three open-model candidates at the precision your GPUs allow. You see accuracy, speed and cost per thousand requests side by side before committing to hardware.
If the private model falls short on a task that matters, the options are a larger model, better retrieval, a fine-tuned adapter via our LLM fine-tuning services, or routing only that task to a hosted API where your rules allow.
How much does private LLM deployment cost in India?
With BtechWaleTech, private LLM deployment setup starts at ₹40,000 (US$600) over 2–4 weeks, and internal applications built on top start at ₹60,000. The ongoing cost is mainly GPU rental or hardware, billed by your provider directly, and it depends on model size, precision, traffic and running hours.
The setup quote grows with the number of models, environments (a test and a production copy doubles some work), security requirements such as isolated networks and audit trails, and applications on top: document search, chat portals or integrations with existing software.
For running cost, the key figure is utilisation. A GPU instance bills whether it is busy or idle. An internal tool used during office hours can be scheduled off overnight and at weekends. A customer-facing service that runs all day needs always-on capacity and possibly spare capacity for peaks. We model two or three scenarios with your provider's current prices so you see the monthly range before launch.
Compare against the hosted alternative honestly. At low or irregular volume, per-token hosted pricing usually wins. At high, steady volume, or when policy leaves no alternative, private deployment makes financial sense. Other providers' quotes vary widely, often because some leave out security hardening, benchmarking or handover.
Private LLM deployment process and timeline
A typical private LLM deployment takes two to four weeks from kickoff to a hardened, tested internal endpoint, with applications on top adding time according to their scope. Most of the calendar goes on benchmarking and security, not on starting the server.
In week one we agree the use cases, collect a test set from your workflows, map data flows and start GPU quota requests in your cloud account. We benchmark candidate models against a hosted baseline on a temporary instance, using only non-sensitive or masked test data at this stage.
In week two we pick the model and precision, size the production instance, build the private network, deploy the inference server in containers and put authentication in front. In week three we load-test with realistic concurrency, set up monitoring and alerts, configure logging and retention, and run prompt-injection and access tests. Week four, where needed, covers application integration, documentation and handover.
Throughout, everything is created in your accounts, with infrastructure written as code so it can be recreated or moved.
- Week 1: use cases, test set, data-flow map, quotas, benchmarks
- Week 2: model choice, network, inference server, authentication
- Week 3: load tests, monitoring, logging, security tests
- Week 4: integration, documentation and handover
Running a private LLM after launch: monitoring, upgrades and ownership
After launch, a private LLM needs the same care as any production server, plus model-specific checks: GPU memory and utilisation, latency, error rates, and a periodic rerun of the quality test set. Someone must own it, and that owner should be in your organisation, with us as support.
Monitoring covers the basics every week: is the GPU saturated at peak hours, are requests queuing, are error rates rising, is disk filling with logs. Alerts go to your team and, during the maintenance period, to us.
New open models appear often, and some will be better or cheaper for your task. Because the test set and deployment are reusable, evaluating a new model is a short exercise rather than a project. Security patches for drivers, containers and the inference server should be applied on a schedule.
Ownership is simple: the cloud account, infrastructure code, container images, configuration and documentation are yours from the start. At handover you receive runbooks for restarting, scaling, rotating keys and swapping models. Two months of free maintenance follow; after that, care starts at ₹8,000/mo a month, only if you want it.
Worked example: a hypothetical lending company in Mumbai
Say a mid-sized non-banking lender in Mumbai wants staff to summarise loan files, draft customer letters and extract fields from income documents, but its policy forbids sending borrower data to any external AI service. This is a hypothetical scenario to show the reasoning, not a client project.
We would start by collecting a hundred masked examples of each task and benchmarking a hosted API baseline against two open models in the 7–14 billion parameter range. Suppose the open models match the baseline on extraction and letter drafting but trail it on long-file summaries.
The plan might then be a 14B-class model at 8-bit on a GPU instance in AWS Mumbai, in a private subnet reachable only from the lender's internal loan system, with vLLM serving an OpenAI-compatible endpoint. Borrower identifiers would be masked in logs, retention set by the compliance team, and every call recorded with the calling service and user. For long summaries, a document-chunking step would feed the model section by section.
The setup would be quoted from ₹40,000, with the staff-facing interface and integration into the loan system quoted separately from ₹60,000. The lender's own counsel would sign off the data flows before launch.
Private LLM deployment checklist
Work through this checklist before and during a private LLM deployment. Items you cannot tick yet become the first tasks in the plan; skipping them is how private deployments end up either leaking data or disappointing users.
Keep the checklist with the handover documents and revisit it whenever you change the model, move clouds or add a new application that calls the endpoint.
- Written reason for going private: policy, volume or isolation
- Test set of real tasks with grading rules
- Model licence reviewed by your counsel
- GPU sizing confirmed by a load test at expected concurrency
- Private network, no public endpoint, restricted outbound traffic
- Authentication and least-privilege roles for every caller
- Logging decided deliberately, identifiers masked, retention set
- Data-flow map including monitoring and backups
- Owner named in your team, with runbooks and alerts
Private LLM deployment for organisations across India
We set up private models remotely for organisations across the country. Financial services firms in Mumbai and policy-sensitive organisations in New Delhi are the most common enquiries, usually driven by contracts or regulators. Manufacturers in Faridabad, Rajkot and Jamshedpur ask about keeping design and process documents off external services.
Hospitals and diagnostic groups in Kozhikode and Warangal want patient-related text processed only on their own infrastructure. Industrial and automotive suppliers in Aurangabad and hospitality groups in Panaji ask about internal assistants that never touch a public AI API.
The method is the same everywhere: test first, size honestly, build in your accounts, lock it down and hand it over with the documents your team needs.