The AI engineer you cannot hire — we cover the seat monthly
RAG, agent orchestration, on-prem inference. These roles are hard to fill: 6–8 weeks from posting to onboarding, and you are still betting on the hire working out. When the roadmap cannot wait, bring in an external team.
Versus hiring
| Hiring in-house | Monthly augmentation | |
|---|---|---|
| Real monthly cost | ≈¥25–60k (salary + statutory contributions + bonus; varies with seniority and city) | Per month / per day / per module — comparable at equivalent seniority; what you save is the cycle and the risk, not the monthly cash |
| Time to start | 6–8 weeks | On board within 3 days |
| Cost of a bad hire | Probation + handover + re-hiring, usually 3 months lost | Stop after the first module if it does not work out |
| When the work ends | The headcount and cost remain | Scale down monthly; no headcount, no payroll liability |
| Invoicing | Payroll — no deductible invoice | VAT invoice, corporate settlement, contract-backed |
What we actually cover
RAG and knowledge bases
Document parsing and chunking, vector store selection and tuning, embedding and rerank deployment, retrieval quality evaluation. Production knowledge bases at 100+ SKU scale.
Agents and orchestration
Multi-agent collaboration, tool calling with fallbacks, Dify or custom orchestration, prompt engineering with regression tests.
On-prem / air-gapped inference
vLLM and Ollama multi-GPU and quantisation, Qwen3 / DeepSeek local inference, offline image and dependency transfer, deployment docs and handover.
Size VRAM and GPUs with the calculator →AI backend engineering
Python / Go / TypeScript services, API gateways, rate limiting and retries, observability — turning a demo into something operable.
When what you lack is not a pair of hands, but someone who can call it
Some blockers are not a capacity problem. Nobody can say which route to take, whether the architecture will hold, or whether the spend is worth it. Another engineer will not fix that, and hiring a credible tech lead takes months.
When you need this
- ·A startup without a technical co-founder, unable to settle on direction
- ·A traditional business going digital or adopting AI, with nobody in-house able to judge vendor proposals
- ·Pre-investment technical due diligence: code quality, architectural debt, real team capability
- ·A system straining after years in production — rebuild, scale out, or re-architect?
What you get
AI adoption roadmap
Which parts of the business genuinely warrant AI and which are wishful thinking; phased roadmap, technology choices, investment scale and acceptance metrics. One-off deliverable you can take straight into planning.
Architecture review
Independent review of an existing system or a proposed design: risk register, priorities and a remediation path — actionable items, not "you should rewrite it".
Technical due diligence
For investors or acquirers: code quality, architectural debt, delivery capability and key-person dependency, delivered as a written report.
Ongoing technical advisor
A set number of days per month on technical decisions, hiring calibration, architecture review and vendor evaluation. No headcount, no day-to-day management.
How it is priced
Priced per engagement or per month, not per man-month — you are buying judgement, not hours. One-off deliverables are fixed-price by scope; ongoing advisory is a monthly retainer for an agreed number of days. We scope the problem first, then quote.
On what basis
20 years in full-stack engineering and technical leadership, with core team members having held technical director / CTO roles at several publicly listed companies. End-to-end delivery across toC, toB and toG platforms, spanning digital publishing and AIGC, e-commerce and retail, O2O, CRM and contact centres, legal tech, robotics and embodied AI, and e-government. The last two years focused on AI application engineering.
How to start
No long-term contract or total price up front. Start with one well-bounded module and let the working relationship prove itself — how requirements get aligned, what the delivery bar is, whether the pace fits. If it works, we talk long term; if not, both sides saved time. Either way it beats screening CVs and waiting out a notice period.
- 01
Tell us where you are blocked and when you need it
- 02
We return an implementation path and schedule
- 03
One well-bounded module, delivered in 1–2 weeks
- 04
Once the collaboration is running smoothly, move to monthly or per-module engagement
Compliance and ownership
- ✓Registered corporate entity: contracts, VAT invoices and corporate settlement
- ✓You work directly with the engineer responsible — no sales layer, no subcontracting
- ✓Source code and IP transfer to you on acceptance, per contract
- ✓NDA supported; air-gapped on-site work available for sensitive projects
When we are not a fit — stated up front
- ·Full-time on-site presence — we work remotely and in scheduled on-site blocks
- ·Generic CRUD backend work unrelated to AI — you will find cheaper elsewhere
- ·Pre-training a foundation model from scratch — we work at the application and engineering layer
- ·Budgets below the per-module minimum — below that line nothing production-grade comes out
Tell us where you are blocked; feasibility within 24 hours
No need for a finished spec. One sentence on the blocker and the deadline is enough.