AI consulting & custom LLM development
Founder-led production LLM engineering — model selection, agent design, and custom LLM features built, hardened, or rescued, and taken past the prototype into production.
Last updated:AI consultancy is the practical integration of large language models into real products — selecting the right model, designing agents, and shipping custom LLM features that hold up in production. Revenant Systems designs, builds, and hardens these features inside existing software, working with whichever hosted or self-hosted model passes evaluation for the workload.
Who is this AI consultancy for?
Revenant Systems' AI consultancy is for CTOs, technical founders, heads of product, and data leads at businesses that already have software, data, and a specific operational problem an LLM could solve. The work is production engineering — building, hardening, and rescuing LLM features — so teams after ChatGPT training, Copilot rollout, or an AI opportunity roadmap are better served by an adoption consultancy.
Which engagement fits where you are?
Three fixed-scope entry points match the three states an AI feature can be in. An idea with no proven implementation starts with the AI feature feasibility sprint. A prototype that now has to face live traffic starts with the AI production-readiness audit. A team that already knows what must be built, integrated, or rescued starts with an engineering engagement, scoped against evidence from one of the two.
Every project price is agreed before work starts, and there is no obligation to continue from a sprint or an audit into implementation.
| Where you are | Entry point | Price |
|---|---|---|
| An idea, not yet proven | AI Feature Feasibility Sprint | From £4,500 |
| A prototype facing production | AI Production-Readiness Audit | From £5,500 |
| A validated feature to pilot | Focused LLM pilot | From £15,000 |
| A feature to build, integrate, or rescue | Production LLM engineering | From £20,000 |
| Not sure which | Technical fit call | Free 20–30 minutes |
Who does the work?
Ian Compton, the founder, leads every engagement and stays directly involved in the engineering. Specialist help is brought in where a project genuinely requires it; nothing is passed to a junior delivery team. The engagement is scoped, built, evaluated, and handed over by the same senior engineer.
A message about a feature goes to Ian, and the usual reply is a free technical fit call — 20 to 30 minutes to establish whether the problem suits the work, whether the data and budget are workable, and which entry point fits.
UK company and contracting entity · UK GDPR-aware delivery · Vendor-neutral model selection · Private and self-hosted deployment options
What do LLM consulting services include?
LLM consulting services cover the work of putting a large language model into a product: choosing the model, designing the prompts, agents, and tool interfaces around it, measuring whether the output is good enough, and running the result once real traffic reaches it. Revenant Systems delivers all of that in one engagement rather than a strategy deck handed to someone else to build.
- A working feature in your codebase
- An evaluation suite your team can re-run
- Guardrails and monitoring already in place
- A handover your engineers can maintain
What is practical AI integration?
Practical AI integration means building LLM features that stay reliable on production data at production volume. Practical AI integration covers model selection, prompt and agent design, evaluation, and the guardrails that stop a wrong answer reaching a user. How much of each a feature needs depends on the stakes: a suggestion someone reviews and an action the system takes alone warrant different evidence.
Revenant Systems selects models per task across Anthropic, OpenAI, Google, and OpenRouter, and runs local models with LMStudio and llama.cpp where UK data residency or cost demands it.
- Model selection and evaluation
- Agent and tool design
- Prompt engineering and guardrails
- Production monitoring
What does a custom LLM build involve?
A custom LLM build pairs a model with your own data, tools, and product rather than adapting an off-the-shelf assistant. The work runs from scoping and model selection through prompt, agent, and tool design, then evaluation, guardrails, and integration into the application — finishing with the production monitoring that keeps the feature reliable once real users arrive. Scope varies by workload: a document-extraction feature is a smaller build than a multi-step operations agent.
Revenant Systems delivers every stage in-house: the same senior engineers handle model selection, agent and tool design, and the application integration, so the build is judged on production behaviour rather than a demo.
- Scope and success criteria
- Model selection against the workload
- Prompt, agent, and tool design
- Evaluation, guardrails, and integration
- Production monitoring and handover
Does a custom LLM feature need RAG or fine-tuning?
Most custom LLM features need retrieval, not fine-tuning. Retrieval-augmented generation (RAG) puts your own documents and records in front of the model at question time, so answers move the moment the underlying data does. Fine-tuning changes how a model behaves, costs more to repeat, and goes stale as the data moves on. Fine-tuning earns its place when a task's format or tone cannot be prompted reliably.
In the agent-driven reporting system Revenant Systems built, every past query and its commentary is embedded and indexed for semantic search, so the agent is handed the most similar prior work — and the failures that were corrected — before it attempts any new SQL.
- Retrieval over your own records, current at question time
- Embeddings and semantic search
- Fine-tuning where format or tone cannot be prompted
- Evaluation to settle which the workload actually needs
Can AI agents drive your existing application?
Yes, through tool exposure: an existing application's search, actions, and workflows are wrapped as typed tools an AI agent can call. Revenant Systems exposes existing applications to LLMs as tools agents can drive, automating multi-step workflows that previously required a person.
- Typed tool and function definitions
- Agent orchestration
- Workflow automation
What do teams use AI integration for?
The recurring use cases are operational rather than novelty: assistants that search and act on existing systems, document workflows that extract and validate structured data, and assessments of whether a workload can run on local models. Each one pairs a model with the tools, guardrails, and evaluation that make it dependable.
| Use case | Why it fits |
|---|---|
| Internal support assistant | Searches and acts on your existing systems through typed tools |
| Quoting assistant | Calls your pricing logic as a tool rather than guessing at it |
| Document review workflow | Extracts and validates structured data with typed, checked outputs |
| Operations assistant | Drives multi-step workflows in an existing application |
| Local-model assessment | Establishes whether data-residency or cost constraints can be met on-premise |
Which models and frameworks does Revenant use?
Revenant Systems works across hosted and local models. Hosted models come from Anthropic, OpenAI, Google, and OpenRouter; local and self-hosted models run through LMStudio and llama.cpp. Application code is typically Python with pydantic-ai, which holds each model call to typed, validated inputs and outputs.
- Anthropic
- OpenAI
- OpenRouter
- Python
- pydantic-ai
- LMStudio
- llama.cpp
AI consultancy, AI platform, or an in-house hire?
Choose by what happens after the feature ships. An AI consultancy suits a defined feature you want built, owned, and handed over — senior engineers, code in your repository, the engagement ending when the feature works. An AI platform suits a workload its product already fits. An in-house hire suits a permanent AI roadmap, once there is enough work to keep an engineer busy.
Revenant Systems has no platform to resell and no affiliation with any model vendor, so the honest recommendation can be any of the three — including telling a client that an existing platform already covers what they need.
| Option | Suits | What you own | Cost shape |
|---|---|---|---|
| AI consultancy | A defined feature that has to survive live traffic | The code, in your repository | Project fee, ending at handover |
| AI platform | A workload the product already fits | Configuration, not code | Licence, ongoing |
| In-house hire | A permanent AI roadmap with steady work | Everything, once hired | Salary, ongoing |
How does Revenant govern its own use of AI?
Against a published statement rather than an assurance. The AI & Responsible AI Statement sets out human oversight of model output, the privacy and data-protection limits on what may reach a model, and the security review that precedes a release. Governance of your own LLM feature, and the obligations it has to satisfy in your sector, is scoped per engagement rather than assumed.
The methodology page describes the firm's own AI practice openly for the same reason: a consultancy that uses AI in delivery should say how.
- Human oversight of model output
- Privacy and data-protection limits
- Security review before release
- Transparency about where AI is used
What's included
- Model selection and evaluation
- Agent and LLM feature design
- Tool / function interfaces for your app
- Guardrails and production monitoring
Stack Anthropic · OpenAI · Google · OpenRouter · Python · pydantic-ai · LMStudio · llama.cpp
Frequently asked questions
Our AI feature works in the demo — will it survive production?
That is exactly the question the AI production-readiness audit answers: a fixed-scope assessment of evaluation coverage, guardrails, prompt and tool design, model choice with its cost and latency profile, and the failure modes that only appear under live traffic — delivered as a severity-ranked report with a hardening roadmap.
Do you build custom LLM features or resell a platform?
Custom builds. The model behind a feature is whichever suits that job — hosted from Anthropic, OpenAI, Google, or OpenRouter, or run locally through LMStudio and llama.cpp — and the feature lands in your product as code your team owns and can change without us. Where an existing platform genuinely fits the need, saying so is part of the advice.
Can our data stay in the UK, or on our own infrastructure?
Yes. Where UK data residency or GDPR obligations rule out sending records to a hosted API, Revenant Systems runs local and self-hosted models through LMStudio and llama.cpp on infrastructure you control, so the data never leaves it. The hosted-or-local decision is made workload by workload — Anthropic, OpenAI, Google, and OpenRouter where a hosted API is acceptable, local models where it isn't.
How long does a custom LLM feature take to build?
Weeks for a single feature; longer for an agent. The two variables are how much the model has to be right about and how much of the surrounding application has to change. A typed extraction or classification feature is a single model call with a schema on each side and clearly defined failure behaviour. A multi-step agent that drives an existing application through tools takes longer, because the tool interfaces, evaluation, and guardrails are most of the work — not the prompt. Fixed-scope audit work is quoted before it starts, and build timelines are scoped against what the audit finds.
How do you stop an LLM feature giving wrong answers?
You can't eliminate model error, but you can measure and contain it. Practical AI integration pairs evaluation — so quality is measured and regressions are caught when models or prompts change — with the guardrails and production monitoring that catch a wrong answer before someone acts on it.
Which model should we use?
The one the evaluation picks. Candidates come from Anthropic, OpenAI, Google, and OpenRouter — judged on quality, cost, and latency for your actual workload — and the choice is revisited as models change, since evaluation coverage makes switching safe.
What does a custom LLM feature cost?
Published starting points: the AI feature feasibility sprint from £4,500, the AI production-readiness audit from £5,500, focused LLM pilots from £15,000, and production LLM engineering from £20,000. Within those, cost is driven by the scope of the feature, the depth of evaluation and guardrails it needs, ongoing model usage, and how much of the surrounding application must change. Model and cloud usage is paid directly by you and stays visible — never bundled into the fee.
Do you use AI to build software yourselves?
Yes — and openly. Nothing reaches a client repository that a person here has not understood and verified, and a senior engineer answers for every change once it is live. The methodology page sets out the practice in full.
Does an AI feature need a data pipeline behind it?
Usually. Model quality rests on data quality — an LLM feature is only as reliable as the data that reaches it, so AI work often pairs with data engineering: the pipelines that feed features dependably sit underneath the model.
Our AI project has stalled — can you rescue it?
Often, yes, and the starting point is evidence rather than a rewrite. The AI production-readiness audit establishes where the feature actually fails and why, then puts the findings in severity order with a route to fixing each one. Sometimes the fix is a redesign; more often it is measurement and guardrails the project never had.
Every engagement follows the same process — see how we work.
The standard this work is held to is published in full: the AI & Responsible AI Statement.
Have an AI feature in mind? Let's talk.
Get in touchYour message is read by the engineer who would scope the work; the reply is a short technical conversation.