Service

AI Agents & Workflows

AI that does actual work, not a chatbot in the corner

Most AI projects fail because they start with the technology and go looking for a problem. I start with the hours your team is losing, work out which of them a machine can genuinely take, and build that. If AI isn't the right tool for the job, I'll tell you and build something simpler instead.

There is an enormous amount of noise about AI right now, and a lot of businesses have been sold something that demos beautifully and then quietly stops being used. The pattern is nearly always the same: it was built to be impressive rather than to fit into how anyone actually works.

What works is narrower and less exciting. Find a specific, repetitive, high-volume piece of work; a machine reads the emails, extracts the information, files it in the right place, chases what's missing, and escalates to a person when it isn't sure. Nobody writes a press release about it, and it saves real money from the first week.

Airwaves is the clearest example. They're a facilities management company that grew fast, and the thing holding them back was inbound jobs: every one arrived by email, got keyed in by hand, and half the time was missing information nobody noticed until an engineer was already on site. The agent I built reads each job, creates the record, and chases the client for anything missing before the job is marked ready. On-site information requests dropped by 98%.

AI agentsRetrieval pipelinesWorkflow automationSystem integrations
Areas I cover

Finding where AI actually pays

Before anything gets built, working out which parts of your operation are worth automating and which aren't. Usually a handful of processes account for most of the wasted hours. Sometimes the answer is that a simple integration would do it and AI is overkill, which is a cheaper and more reliable outcome for you.

Document & email processing

The highest-value automation for most businesses, because it's where the manual typing lives. Reading inbound emails, invoices, orders, forms, and PDFs, pulling out the information that matters, and putting it into your systems correctly, including the ones with no API worth speaking of.

Tool-using agents

Agents that don't just answer questions but do things: create records, send emails, update your CRM, chase missing information, and hand off to a person when they hit something they shouldn't decide alone. Connected to the systems you already run rather than a separate place to check.

Retrieval over your own knowledge

Answers grounded in your documents, your policies, and your history, rather than whatever the model absorbed from the internet. This is what stops an AI confidently inventing an answer, and it's the difference between something your team trusts and something they quietly stop using.

Guardrails & human-in-the-loop

Deciding upfront what the system is allowed to do on its own, what needs a person to approve it, and what it must never touch. Anything irreversible or expensive gets a human in front of it. This gets designed in from the start, not bolted on after something goes wrong.

Evaluation & monitoring

Measuring whether it's actually working, on your real data, with the failures visible rather than buried. Models and vendors change underneath you, so something that worked in March can drift by September. Without monitoring you find out from a customer, which is the expensive way.

How it works
01

Find the hours

I sit with you and your team and work out where the time genuinely goes: the repetitive, manual, error-prone work that's eating the week. This is also where I'll tell you which parts AI shouldn't touch.

02

Build with guardrails

I build the agents, workflows, and integrations, with the limits agreed in advance and a person in the loop wherever the decision warrants one. It goes live on a narrow slice first so you can see it working on real work before it's trusted with everything.

03

Measure and expand

We check the numbers against what you were doing before, tune what's live, and widen the scope where it's earning its keep. If a piece of it isn't paying for itself, I'd rather turn it off than defend it.

Proof
30 hrs
Saved per week - Cartwright Hands
4 hrs
Saved per person, per day - Knights Events
98%
Fewer on-site info requests - Airwaves

These are real figures from real clients, not modelled averages or vendor benchmarks. Airwaves' inbound job processing is now fully automated: administrative staff no longer spend their day on data entry, jobs reach engineers faster, and engineers turn up with the information they need instead of ringing the client from the car park.

Read the Airwaves case study →
Frequently asked questions
What's the difference between an AI agent and a chatbot?+

A chatbot answers questions. An agent does the work. The agent I built for Airwaves reads inbound job emails, extracts the details, creates the job record in their system, and emails the client directly to chase anything missing, without anyone asking it to. Nobody sits in front of it typing. That's the distinction that matters commercially: one is a better search box, the other removes a job from someone's day.

Is this just automation with extra steps?+

Sometimes, and when it is, I'll say so and build the plain automation instead. It's cheaper and it fails less. AI earns its place specifically where the input is messy and unpredictable, like free-text emails from a hundred different clients who all format things differently. Traditional automation needs the input to be consistent. Where yours already is, you don't need a model in the middle of it.

Where does AI genuinely make sense, and where doesn't it?+

It makes sense on high-volume, repetitive work where the input varies but the outcome is well-defined: reading documents, triaging inboxes, extracting data, drafting routine replies. It makes far less sense for judgement calls with real consequences, anything needing guaranteed accuracy on every single item, or work you only do occasionally. The occasional stuff isn't worth automating regardless of the technology.

What happens to our data? Does it leave the business?+

That depends on the design, and it's a decision you should make deliberately rather than discover afterwards. There are setups where your data never leaves infrastructure you control, and setups using third-party model providers with contractual guarantees about training and retention. I'll lay out the options with the actual trade-offs, including cost and capability, before anything is built. If you're in a regulated sector, we start from your constraints, not from what's convenient.

What happens when it gets something wrong?+

It will, occasionally, which is precisely why the guardrails get designed first. The system knows what it's allowed to do alone and what needs a person to sign off, and anything irreversible or expensive sits behind a human. When it isn't confident, it escalates instead of guessing. The right question isn't whether it will ever be wrong, but what happens when it is, and your existing manual process has an error rate too.

Do we need our own AI model?+

Almost certainly not. Training your own model is expensive, slow, and rarely the reason a project succeeds. What actually makes the difference is giving a good existing model the right context from your own systems and documents, and wiring it properly into your processes. If anyone is proposing training something bespoke for a standard business automation problem, ask them what it buys you.

How do you know whether it's actually working?+

By measuring it against what you were doing before. That means agreeing upfront what we're counting, whether that's hours, error rates, or turnaround time, and then checking. The Airwaves 98% figure exists because we knew how often engineers were requesting information on site beforehand. Anything running without that measurement in place is asking you to take its usefulness on faith.

Will it work with the systems we already use?+

Usually, yes, and that's the point; an AI that lives in a separate window is another thing to check rather than a saving. If your systems have APIs, connecting is straightforward. If they don't, there are normally still routes in, though I'll be honest about which ones are more fragile and need watching.

What does this cost, and how soon does it pay back?+

There's a build cost and an ongoing running cost, and you should see both before you commit. What drives the build is how many systems it touches and how messy the input is; the running cost scales with volume and is usually modest against the hours saved. Payback on a well-chosen first automation is typically fast, because we deliberately start with the process wasting the most time. If the numbers don't work, that's a reason not to build it and I'll tell you.

Can you start small rather than automating everything?+

That's how I'd prefer to do it. One narrow process, live and measured, so you can see whether it works on your real data before committing further. It's lower risk for you and it settles the argument with the numbers rather than a demo. Once one is genuinely earning its keep, the case for the next one makes itself.

Contact

Let's build something.