Build

    Agentic AI, RAG and MCP

    Retrieval grounded in your own knowledge, and agents that can act in your systems — built with the permissions, evaluation and audit trail that make them safe to switch on. Including an honest answer about where an agent should not be given the keys.

    Duration

    6–14 weeks

    Built on

    ISO/IEC 42001 · ISO/IEC 27001:2022 · OWASP Top 10 for LLM Applications · EU AI Act

    Indicative price

    €25,000–65,000 per engagement

    Who this is for

    • CIO or CTO

      Is being asked for agents and is accountable for what they are allowed to touch.

    • Head of operations

      Has a workflow spanning four systems and too many pairs of hands.

    • Knowledge or information manager

      Has the documents, and no way for anyone to find the answer inside them.

    • CISO

      Sees a class of system that reads everything and acts on some of it.

    An agent is a permission problem wearing a language interface

    Two distinct things are sold as one. Retrieval — giving a model access to your documents and data so its answers are grounded in what you actually hold. And agency — allowing a model to take actions in your systems: raising a ticket, sending a message, updating a record. The first is mostly an information architecture problem. The second is an access control problem, and the failure modes are not comparable.

    Retrieval fails quietly. The corpus holds the superseded policy alongside the current one and nothing marks which is which. Permissions were not inherited, so the model surfaces an HR document to the wrong person, politely. Chunking separated the table from its caption. None of these produce an error. They produce a confident, wrong, well-formatted answer, which is worse than an error.

    Agency fails loudly. An agent holding a credential has that credential's full reach, and it will find paths through it nobody enumerated. The controls that matter are ordinary: least privilege, a separate identity per agent, human confirmation before anything irreversible, spend and rate limits, and a log good enough to reconstruct what happened. They are unglamorous, and they are the whole job.

    MCP has made connecting models to systems straightforward, which is exactly why the permission question matters more than it used to. Every connected server is a new path into a system, and 'the model can reach it' and 'the model should be allowed to reach it' are different statements that are very easy to conflate during a demo.

    So we build in that order. Retrieval first, grounded and permission-aware, because a large share of what people ask agents for is actually a retrieval problem. Then agency, narrowly, with irreversible actions behind a human. Then we widen it on evidence from the logs, rather than on enthusiasm.

    How we do it

    1. 01

      Use case and boundary

      1 week

      What the system should do, and specifically what it must never do. The boundary is a design input, not a policy written afterwards.

    2. 02

      Knowledge and permissions

      2–3 weeks

      The corpus: what is in it, what is superseded, and who may see what. Retrieval respects the permissions of the person asking — the single most common gap we find in systems already built.

    3. 03

      Retrieval build and evaluation

      2–3 weeks

      Grounded answers with citations back to source, evaluated against a question set with known answers. Refusal behaviour tested as carefully as correct answers: a system that will not say 'I do not know' is a liability.

    4. 04

      Agency design

      1–2 weeks

      Which actions, under which identity, with what limits, and which require a human. Separate identity per agent, least privilege, irreversible actions confirmed by a person.

    5. 05

      Integration and guardrails

      2–3 weeks

      MCP servers or direct integrations, each scoped individually. Prompt injection handling, rate and spend limits, and a log that reconstructs the sequence behind any action.

    6. 06

      Pilot, monitor, widen

      2–3 weeks

      A narrow live pilot, monitored. Scope widens on what the logs show rather than on how the demo felt.

    Named artefacts

    What you receive

    • Grounded retrieval over your corpus, with citations to source
    • Permission-aware retrieval, inheriting your existing access rights
    • Evaluation set with known answers, including refusal cases
    • Agent action inventory — what it may do, under which identity, within what limits
    • Human confirmation design for irreversible actions
    • Integration or MCP server configuration, scoped per system
    • Prompt injection and abuse handling, tested before go-live
    • Audit log sufficient to reconstruct any action taken
    • Spend and rate controls with alerting
    • Runbook, including how to stop it

    What we need from you

    • Somebody who can say what is authoritative in the corpus. Retrieval over contradictory documents produces confident contradictions.
    • Your access model, and a willingness to fix it where it is wrong. An agent inherits whatever mess it is handed.
    • A decision on which actions require a human. We will recommend; the decision is yours to own.
    • Test questions from the people who will actually use it, including the awkward ones.

    What changes

    1. 01Answers grounded in your own material, with a citation the user can check.
    2. 02Nobody sees through retrieval what they could not see directly.
    3. 03Agents act under their own identity, within limits, with irreversible steps behind a person.
    4. 04Every action is reconstructible from the log.
    5. 05Scope widens on evidence, and there is a documented way to stop it.

    What it costs

    €25,000–65,000 per engagement

    All prices exclude VAT.

    Questions

    Do we need an agent, or is this a retrieval problem?

    More often than not it is retrieval. 'Let the team ask questions of our documents' needs no agency at all, and adding it multiplies the risk for no benefit. We are happy to build the smaller thing when the smaller thing is the answer.

    Is our data used to train somebody's model?

    That depends entirely on the provider and the contract, and it is a question to settle in writing before anything is built. We will tell you what each option means for your data and, where the requirement is that nothing leaves a boundary, design for that instead.

    What about prompt injection?

    A real class of attack, handled with architecture rather than instructions: treat retrieved content as untrusted, keep privileges minimal and separated, require confirmation for consequential actions, and log enough to detect it. We test for it before go-live.

    Can an agent work inside our existing systems?

    Usually, through MCP or direct integration. The constraint is rarely technical — it is what the identity behind it is permitted to do, which your security function should decide deliberately rather than inherit from a service account.

    How do we know it is not quietly getting worse?

    The evaluation set runs on every change, and the logs are monitored for refusals, escalations and abandoned interactions. Both matter: quality regressions show up in the test set, usefulness problems show up in behaviour.

    Leave with your top three risks documented

    Thirty minutes with a senior practitioner. No slideware, no sales engineer.