Alien Private AI Server
Useful AI does not always need your sensitive data in a public consumer account
Internal documents, support history, procedures, code and client records can make an AI assistant more useful. They also make the data path more important.
Alien Private AI Server keeps selected model inference and document retrieval on infrastructure the customer controls. It is designed for clearly defined workloads where privacy, custody or predictable local access justify the hardware and operating responsibility.
It keeps the right workloads in-house without claiming every workload belongs there.
What "private AI" needs to cover
Running model weights locally is only one part of the path. A useful deployment may also include:
- user interface
- authentication and permissions
- model runtime
- model files and licences
- prompt and response logs
- document ingestion
- text extraction and chunking
- embedding model
- vector or search database
- backup and recovery
- update and evaluation process
- optional connectors or external search
If any component calls an external service, that flow must be documented. A local model behind a cloud transcription plug-in is not an entirely local workload.
Possible capabilities
Depending on the approved use case and hardware, the service can provide:
- local chat and drafting
- private search across approved document collections
- retrieval-augmented generation (RAG)
- summarisation and classification
- internal procedure and support assistance
- local embeddings and vector search
- role-based access to collections
- source citations back to internal documents
- controlled model and application updates
- usage and quality evaluation
The design begins with a real task and acceptance test. "Install an LLM" is not a business outcome.
How local inference changes the data flow
A local runtime can process prompts on the host rather than sending them to the model provider. Ollama states that it does not see prompts or data when the software runs locally and offers a setting to disable its cloud features (Ollama). llama.cpp is another project designed to run LLM inference across local and cloud hardware (llama.cpp).
Those statements describe the runtime. Alien IT separately checks the interface, plug-ins, connectors, telemetry, operating system, monitoring and remote support.
What private RAG does
RAG retrieves relevant passages from an approved document collection and supplies them to the model with the user's question. The model can answer with links or citations to those passages.
This can make an internal assistant more grounded and current than relying on model memory alone. It does not guarantee truth.
The system still needs:
- document-level access controls
- a deletion process that reaches the search index
- protection against malicious or misleading documents
- version and freshness signals
- evaluation questions with expected evidence
- limits on what the assistant may recommend or automate
If the search index ignores the source system's permissions, RAG can expose information to the wrong employee more efficiently than ordinary search. Permission design comes first.
What it can reduce
A correctly scoped local deployment can reduce:
- routine upload of prompts and source documents to an external consumer AI service
- dependence on a provider's changing consumer settings
- internet dependence for supported local tasks
- uncontrolled mixing of personal and business AI accounts
- external custody of the local prompt and retrieval history
It can also give a business a controlled place to test models and data collections before broader use.
What it cannot guarantee
An Alien Private AI Server does not guarantee:
- accurate, current or unbiased answers
- suitability for legal, medical, financial or safety-critical decisions
- that every model licence permits every commercial use
- security without patches, access control and monitoring
- confidentiality if users can export results or source documents
- deletion from backups immediately
- quality equal to the largest hosted models
- lower total cost for every usage level
- operation during hardware failure without a recovery design
Local hosting changes custody. It does not make an inaccurate answer safe, it does not secure the host automatically, and it does not replace backups.
When hosted business AI may be the better choice
Major providers offer business products with different data treatment from consumer products. OpenAI says API, ChatGPT Business and Enterprise content is not used for training by default (OpenAI). Microsoft says prompts, responses and Graph data in Microsoft 365 Copilot are not used to train foundation models and stores interactions under Microsoft 365 controls (Microsoft).
A hosted business service may be preferable when:
- frontier model quality is essential
- usage is intermittent and hardware would sit idle
- the provider's contract and controls meet the need
- fast product improvements matter more than local operation
- a small team cannot safely maintain AI infrastructure
- broad integrations are core to the workflow
The right comparison is a specific hosted product and contract against a specific local architecture, not "cloud AI" against "private AI" as slogans.
Hardware and performance
Model size, quantisation (a compressed copy of the model), context length (how much text it can consider at once), concurrent users and output speed determine hardware needs. More memory can run a larger model; a capable GPU can improve throughput; storage must hold model versions and indexes. Power, heat, noise and replacement time belong in the design.
We benchmark the proposed workload with representative documents and questions. A model that produces an impressive demo but cannot meet peak use, cite sources or handle the organisation's language is not ready.
Security and operations
A production deployment requires:
- named administrators and users
- least-privilege access to document collections
- encrypted administration and remote access
- restricted network exposure
- how long prompts and logs are kept
- where each model and software component came from
- update and vulnerability process
- backup of configuration, indexes and required source data
- restore test
- quality and misuse monitoring
- an incident and shutdown path
Do not expose the model server directly to the public internet for convenience. Remote use should pass through an authenticated, supported access layer.
Engagement stages
1. Use-case and data assessment
Choose one bounded task. Classify the data, users, expected answer quality, prohibited uses, latency and availability needs.
2. Hosted-versus-local comparison
Compare approved hosted business services, local models and a hybrid design against custody, quality, cost, skill and exit criteria.
3. Proof of value
Run representative questions against an approved sample. Score source quality, answer correctness, refusal behaviour, speed and hardware use.
4. Controlled build
Configure identity, model runtime, retrieval, logging, backups and network access. Record every external connection.
5. Acceptance and red-team checks
Test permission boundaries, prompt injection through documents, unsupported questions, deletion propagation, outage and restore.
6. Handover or management
Agree who owns models, patches, users, data ingestion, evaluation and recovery. A private server with no operating owner will become an insecure old server.
Who it suits
This service may suit organisations with a repeatable internal knowledge task, sensitive documents, enough sustained use to justify hardware, and a named person or managed agreement responsible for operation.
It may not suit occasional public-information questions, work that requires the best hosted model at all times, teams without a defined data set or evaluation method, or high-stakes automated decisions.
Frequently asked questions
Does local AI send nothing to the internet?
It can be configured for local-only inference, but updates, model downloads, external search, plug-ins or remote support may create separate connections. The final data-flow register is the answer for the deployed system.
Can it search our company documents?
Yes, where the document formats, permissions and quality are suitable. Ingestion and deletion rules must be agreed first.
Will it be as good as ChatGPT, Claude or Gemini?
That depends on the model and task. Smaller local models can be excellent for narrow work and weaker for complex reasoning or broad knowledge. Alien IT benchmarks the actual use case.
Is it cheaper?
Sometimes at sustained utilisation. Hardware, power, maintenance, support, downtime and model evaluation must be compared with hosted fees.
Can it train our own model?
The base service focuses on inference and retrieval. Fine-tuning or training is a separate project with additional data, licensing, evaluation and compute requirements.
Does it make AI safe for confidential information?
It can remove a major external data path. Confidentiality still depends on users, permissions, endpoint security, logs, backups, physical access and any integrated service.
Keep selected intelligence in-house
Each workload you keep on infrastructure you control is one more set of documents that never needed uploading to a consumer AI service.
Ask Alien IT to test one private-AI use case on your own infrastructure. Call 02 9707 0999 or use the contact page.