Services
Everything we do, and who it's for.
Four core practices plus the follow-on work clients most often need. Every engagement ends with you owning the source, the documentation, and the ability to run it without us.
Core practices
Where most engagements start
Agent Skills & SKILL.md automation
For: teams whose expertise lives in people's heads and pasted prompts.
Encode repeatable workflows as versioned Agent Skills — a folder with a SKILL.md, reference documents, and executable scripts that an AI agent loads only when the task calls for it. We handle authoring, progressive-disclosure design, script bundling, testing, org-wide skill libraries, versioning, and governance.
Extension & integration support
For: teams whose AI can't see the systems the work actually lives in.
IDE extensions (VS Code, JetBrains), browser extensions, Slack and Teams apps, MCP servers exposing your internal tools, and connectors to Jira, Salesforce, HubSpot, Notion, GitHub, Zendesk, and Google Workspace. OAuth, scoped permissions, rate limiting, and audit logging designed in from the start.
Full detail →API & backend development
For: teams with a promising prototype and no path to production.
Streaming endpoints, tool-use and agent loops, retrieval pipelines, structured outputs, prompt caching, cost controls, background jobs, webhooks, evaluation harnesses, observability, and CI/CD. Python, TypeScript, and Go, deployed to Cloudflare, AWS, GCP, Azure, or on-premises.
Full detail →AI coaching for non-technical teams
For: everyone in the building who isn't an engineer.
Plain-English, hands-on sessions using participants' real work. Covers daily-life use, role-specific workflows, what AI gets wrong, what data must never go in, and how to verify an answer. Delivered as workshops, a multi-week programme, executive briefings, or 1:1 coaching.
Full detail →Also in demand
The work that comes next
These are the requests that follow a first successful project — and increasingly, the reason clients call in the first place.
AI readiness, policy & governance
A short, readable AI use policy people will actually follow. An approved-tools list with reasons. Data classification rules covering what may and may not be sent to a third-party model. A risk register, an incident path, and a record of decisions — the artefacts your legal, compliance, insurance, and enterprise-procurement reviews are now asking for.
Typical output: policy document, tool register, data-handling matrix, staff one-pager, leadership briefing.
Document & data extraction pipelines
Invoices, contracts, purchase orders, claim forms, CVs, inspection reports, and email attachments converted into validated structured data — with confidence scoring and a human review queue for anything uncertain. Usually the clearest measurable ROI available, because you can count the hours it replaces.
Typical output: ingestion pipeline, schema, validation rules, review UI, accuracy report.
Internal knowledge assistants (RAG done properly)
A chat interface over your handbook, policies, wiki, tickets, and past project files that cites its sources and admits when it doesn't know. Built with your existing permission model respected at retrieval time, so people never see documents they shouldn't — the step most quick builds skip.
Typical output: indexing pipeline, permission-aware retrieval, cited chat UI, eval set, refresh schedule.
Cost, latency & usage optimisation
For teams whose AI bill or response times went sideways. We audit real traffic and apply prompt caching, model routing (frontier models only where they earn it), batch processing for non-urgent work, effort tuning, and context management — then add per-team budgets, alerts, and dashboards so it stays fixed.
Typical output: usage audit, optimisation PRs, spend dashboard, budget alerts, before/after benchmark.
Evaluation & quality harnesses
You cannot safely change a prompt, a model, or a retrieval strategy without a way to measure whether things got worse. We build golden datasets from your real cases, scoring (deterministic checks where possible, model-graded where not), regression gates in CI, and a dashboard non-engineers can read.
Typical output: eval dataset, scoring harness, CI integration, quality dashboard, review process.
Fractional AI lead / advisory retainer
Senior capacity by the day: reviewing vendor claims and contracts, unblocking your engineers, setting standards and code review for AI work, running an internal community of practice, and giving leadership a straight read on what's feasible this quarter versus what's a keynote fantasy.
Typical output: standing sessions, architecture reviews, vendor assessments, roadmap input.
Workflow & agent automation
Multi-step processes where an agent plans, calls tools, and checks its own work — triage, research, reconciliation, reporting, QA passes. Built with explicit approval gates on anything irreversible, budget ceilings, and full traces, because an unsupervised agent with write access is a liability, not a feature.
Typical output: agent definition, tool surface, approval gates, budget caps, trace logging.
Migration & modernisation
Moving off a deprecated model, a discontinued API parameter, or a vendor you've outgrown — without a regression cliff. Includes a compatibility audit, prompt re-tuning for the new model's behaviour (they are genuinely different), a token and cost re-baseline, and a staged cutover behind evals.
Typical output: audit report, migration PRs, re-tuned prompts, before/after eval results.
Engagement shapes
Four ways to work together
| Shape | Best when | Typical length | You get |
|---|---|---|---|
| Audit & roadmap | You're not sure where AI would actually help, or you need something written for leadership. | 1–2 weeks | Prioritised opportunity list with effort and value estimates, risk notes, and a sequenced plan. |
| Fixed-scope pilot | You know the workflow; you need it built properly the first time. | 2–6 weeks | Working software, tests, docs, runbook, and a handover session. Fixed price. |
| Build partnership | Ongoing delivery across several workflows, or embedded alongside your team. | 3+ months | Sprint-based delivery, shared backlog, code review, and standards your team keeps. |
| Coaching & enablement | The tools are fine — the adoption isn't. | 1 day – 8 weeks | Workshops or a rolling programme, recorded sessions, prompt library, and internal champions. |
Not sure which fits? That's exactly what the free scoping call is for — and a recommendation of "none of these, do X instead" is a perfectly normal outcome.
Start with one workflow.
Bring the task your team complains about most. Thirty minutes, no obligation, written recommendation either way.