There are two ways to put a generative AI feature into a product. You can assemble the platform yourself — gateway, orchestration, retrieval, guardrails, evaluation, observability — which is six months of work before the first feature ships. Or you can rent that platform and spend those six months on the feature. Most organisations should do the second, and a surprising number do the first because nobody framed it as a choice.
This guide is about that choice. What a managed generative AI platform actually provides, what it conspicuously does not, how the economics compare once you account for the parts nobody budgets for, and how to adopt one without ending up unable to leave.
What you will learn
- What sits between a model API and a production feature, and who builds it
- The three tiers of the market and what each is genuinely for
- The build-versus-buy calculus, including the costs usually omitted
- What remains yours regardless of what you buy
- How to evaluate providers on the things that matter after the demonstration
- Keeping the exit open without sacrificing the benefit
- The gap a platform fills
- The three tiers of the market
- What a managed platform provides
- What it does not provide
- The economics, honestly
- What remains yours
- Data handling and compliance
- Evaluating providers
- Cost control and predictability
- Reliability and the dependency
- Lock-in and the exit
- The hybrid position
- Adopting one without regret
- When to build instead
- Twelve mistakes
- A worked example: two teams, same feature
- A capability checklist for the evaluation
- Frequently asked questions
1. The gap a platform fills
A model API gives you one thing: text in, text out. A production feature needs considerably more, and enumerating it is the clearest way to understand what a platform is selling.
| Capability | Why the feature needs it |
|---|---|
| Prompt management | Versioned, reviewable, rollback-able — not a string in code |
| Retrieval | Grounding answers in your data, which is the largest quality lever |
| Model routing | Cheap model for most requests, capable model when needed, fallback on failure |
| Caching | Repeated questions answered without inference |
| Guardrails | Input screening and output validation that can reject |
| Evaluation | Knowing whether a change improved anything |
| Observability | Traces, tokens, cost, and the ability to explain a past answer |
| Budgets and quotas | So one team's loop does not consume the month |
Every one of these is required whether you build it or rent it. That is the whole argument: the work does not disappear, so the question is only who does it, and whether doing it yourself is a differentiator or a tax.
2. The three tiers of the market
The category is confusingly named, and three quite different things get sold under it.
Model providers. The models themselves, accessed by API. They handle inference, availability and scaling. Everything above is yours. Cheapest per token and the largest amount of remaining work.
Platform layers. Products sitting between your application and the models, providing routing, caching, prompt management, evaluation, guardrails and observability. You still write the feature; you stop writing the plumbing. This is the tier most people mean by GenAI-as-a-Service, and it is where the build-versus-buy question actually bites.
Applied products. Finished capabilities — a support assistant, a document extractor, a meeting summariser — that you configure rather than build. Fastest to value and least differentiated, since your competitors can buy the same thing.
These are not exclusive, and most organisations end up using all three: an applied product for a commodity capability, a platform layer for the features that matter, and direct model access for anything unusual. Deciding which tier each use case belongs to is more useful than choosing one tier for everything.
3. What a managed platform provides
Beyond the capability list, four things are worth naming because they are what you are actually paying for.
Time. A feature shipping in three weeks rather than five months. For most products this dominates every other consideration, because the value of arriving earlier exceeds the difference in running cost by a wide margin.
Operational absorption. Provider outages, model deprecations, rate limit changes and the constant churn of the underlying ecosystem become someone else's problem — or at least someone else's first response. This is a real and recurring cost that build calculations reliably omit.
Accumulated judgement. A mature platform encodes hundreds of decisions somebody already got wrong once: how to chunk documents, when to rerank, how to structure retries, what to cache and for how long. Rebuilding that knowledge takes iterations you can skip.
Cost engineering you would not do. Semantic caching, routing to cheaper models, prompt compression and batch processing are all things a platform does by default and a first-party implementation typically adds in year two, after the invoice becomes a topic.
4. What it does not provide
Being precise here prevents the most common disappointment.
- Your data pipeline. Getting your documents into a state worth retrieving from — extraction, cleaning, permissions, freshness — is your work regardless. It is frequently the largest single effort in the project.
- Your evaluation criteria. A platform runs the harness; only you can say what a good answer looks like for your domain.
- Your domain knowledge. Which questions matter, what the correct answers are, and where the edge cases lie.
- Your interface. How the capability appears in your product, what happens when it is unsure, and how a user corrects it.
- Accountability. If the feature says something wrong to a customer, that is yours. No agreement transfers it.
- The decision about what to build. Platforms accelerate execution and do nothing about direction.
The pattern: platforms absorb the generic and leave the specific. That is the correct division, and it is also why buying does not shorten the project as much as the demonstration suggests — the specific parts were always the long ones.
5. The economics, honestly
Build calculations are almost always wrong in the same direction, because they count the visible cost and omit the recurring one.
| Cost | Build | Buy |
|---|---|---|
| Initial platform work | Substantial engineering effort before the first feature | Effectively none |
| Inference | Provider rates directly | Provider rates plus a margin, offset by better routing and caching |
| Ongoing maintenance | Continuous — the ecosystem moves fast | Included |
| Incident response | Yours, including at night | Shared, with support |
| Keeping up with capability | Deliberate work each quarter | Arrives |
| Opportunity cost | The features not built during the platform work | None |
The line that decides most cases is the last one. Engineering time spent building a gateway is engineering time not spent on the product, and unless your platform layer is somehow better than what you could rent, that time bought nothing a customer will notice.
The genuine argument for building arrives at scale: at high sustained volume, a percentage margin on inference becomes a large absolute number, and internalising the platform starts to pay. That crossover is real and it is much further out than teams assume — and by the time you reach it you will know your requirements well enough to build the right thing, which you do not at the start.
6. What remains yours
Regardless of what you buy, five things stay with you, and they are the ones that determine whether the feature is good.
The knowledge. Your documents, their structure, their permissions and their freshness. Retrieval quality caps answer quality, and retrieval quality is mostly a function of how well your content is organised — which no platform can fix for you.
The evaluation set. Representative questions with known-good answers, built from real traffic. This is the single most valuable asset in an AI feature, it takes months of real usage to accumulate, and it is portable between platforms.
The product decisions. What the feature does when it is uncertain. Whether a human reviews. How a user corrects a wrong answer. These determine whether people trust it.
The domain rules. What may be said, to whom, under what conditions. Generic safety filtering is bought; your specific policy is not.
The measurement of value. Whether the feature is worth its cost, measured in something the business recognises.
A useful way to think about it: buy the machinery, own the corpus and the evaluation set. Those two are what make the feature yours and what you take with you if you change providers.
7. Data handling and compliance
Sending your content and your customers' questions to a third party is a data transfer, and it needs the same treatment as any other.
The questions with documented answers, not assurances:
- Retention. How long are prompts and responses stored, and can that be reduced or eliminated?
- Training use. Is your data used to improve their models, and is opting out contractual or a setting?
- Processing location. Which jurisdictions, including for support access and logging?
- Subprocessors. Which model providers sit behind the platform, and are you notified when that changes?
- Deletion. Can you require deletion, and how is it evidenced?
- Certification. What independent assurance exists, and when was it last assessed?
The subprocessor question deserves emphasis, because a platform layer typically routes to several model providers. Your data protection assessment must cover the chain, not just the first link, and a platform unwilling to disclose it is not viable for regulated data regardless of its capabilities.
On your side: minimise what you send, redact identifiers that add nothing, log what was sent, and classify use cases so that the most sensitive ones can be routed differently or kept internal.
8. Evaluating providers
Demonstrations are uniformly impressive, which makes them useless for discrimination. These questions are not.
| Question | What a weak answer reveals |
|---|---|
| How do we evaluate quality on our own data? | No evaluation capability means no way to know if changes help |
| What happens when your upstream model provider has an outage? | No fallback means their incident is your incident |
| How do we export our prompts, evaluations and retrieval configuration? | Vague answers mean the exit is closed |
| Can we bring our own model keys? | Determines whether you can negotiate inference separately |
| How is cost attributed per feature and per tenant? | Aggregate billing means you cannot optimise |
| What does the audit trail contain? | Determines whether you can explain a past answer |
| How do guardrails behave — reject, or log? | Logging-only guardrails are metrics, not controls |
| What is the deprecation notice period for models? | Short notice means unplanned work with a deadline |
Run a genuine pilot rather than a demonstration: your data, your questions, your evaluation set, measured. Two weeks of that tells you more than any comparison document, and it also builds the evaluation set you will need regardless of the outcome.
9. Cost control and predictability
The most common unpleasant surprise with these platforms is not the rate; it is the variance. A feature costing a predictable amount per month is manageable; one that costs ten times more in a month when a customer bulk-imported data is a finance conversation.
The controls to insist on and configure from the first day:
- Per-tenant and per-feature budgets that actually reject, not just alert.
- Cost visible per request in your own telemetry, not only in the provider's dashboard.
- Caching enabled and measured. A cache hit rate you can see is a cost lever you can pull.
- Routing configured deliberately. Most requests do not need the most capable model, and defaulting everything to it is the single largest avoidable cost.
- Bounded loops. Any agentic behaviour needs a hard ceiling on steps and tokens.
- Output length limits matched to what your interface actually shows.
Measure cost per resolved outcome rather than per call. A cheap answer that produces a support ticket is the most expensive result available, and per-call metrics make it look like a success.
10. Reliability and the dependency
Adding a platform adds a dependency, and it sits in your critical path. Two consequences.
Their availability becomes part of yours. Ask for the actual commitment and the historical record, and design as though it will be breached, because eventually it will be. A platform that itself routes across several model providers is more resilient than one tied to a single upstream — which is a genuine argument for the platform layer over direct model access.
You need a defined degradation path. Primary route, fallback model, cached answer, honest decline. Build this before launch and rehearse it, because designing a fallback during an incident produces a worse fallback. An interface that says "this is unavailable right now" is considerably better than one that hangs for thirty seconds.
Also worth confirming: whether their rate limits are shared across their customers or reserved for you, because a noisy neighbour on a shared quota is a failure mode you cannot see coming.
11. Lock-in and the exit
Lock-in is real and is frequently overstated. What matters is which parts are actually hard to move.
Portable: your prompts, your evaluation set, your source documents, your product logic. These are the valuable parts, and they move easily if you keep them in your own repository rather than only in the platform.
Moderately sticky: retrieval configuration, guardrail rules, routing policy. Conceptually portable, practically a few weeks of reconfiguration.
Genuinely sticky: deep integrations with platform-specific abstractions, and any workflow logic expressed in their proprietary format rather than in your code.
The practices that keep the exit open cost little: define your own interface for AI capability so application code never calls the platform directly; keep prompts and evaluation sets in your repository as the source of truth; store source documents in your own systems and let the platform index rather than own them; and record enough per-request detail in your own telemetry that you are not dependent on their dashboard for history.
Do those four and switching becomes a project rather than a rewrite, which is the realistic goal. Refusing to use anything platform-specific in the name of portability means using none of the value you paid for.
12. The hybrid position
Most mature organisations end up somewhere between build and buy, deliberately.
The pattern that recurs: buy the platform, own the knowledge and evaluation, and keep one capability internal — usually the one that is either most sensitive or most differentiating. That gives you the speed of the platform, retains the assets that matter, and preserves the internal skill to build more if the economics change.
A second common hybrid is by data classification: public and general-purpose features on the managed platform, and anything touching the most sensitive data handled by a self-hosted model within your own boundary. This is more work and it resolves compliance questions that would otherwise block the whole programme.
What does not work well is running two full platforms in parallel without a clear rule for which is used when. That produces duplicated evaluation, inconsistent behaviour and twice the operational surface, for the benefit of avoiding a decision.
13. Adopting one without regret
- Settle the data position first. What may be sent, under what agreement, with what retention. Before a pilot, not after.
- Pick one narrow use case with a measurable outcome and a human in the loop. Not the most strategically important one.
- Build the evaluation set from real traffic during the pilot. This is the asset you keep regardless of what you decide.
- Define your own interface so application code depends on your abstraction rather than theirs.
- Instrument cost and quality in your own telemetry from the first request.
- Configure budgets, routing and caching deliberately rather than accepting defaults.
- Rehearse the degradation path before launch.
- Review after a quarter against the outcome you named, not against how impressive it feels.
14. When to build instead
Building the platform yourself is the right answer in identifiable circumstances:
- Sustained very high volume, where the margin on inference exceeds the cost of the engineering to remove it.
- Regulatory constraints that no provider can satisfy — data that genuinely cannot leave your boundary, in a jurisdiction where no acceptable region exists.
- The platform is your product. If you sell AI capability, the platform is not overhead.
- Requirements no platform serves, such as unusual model architectures or an execution environment nobody supports.
- You already have the platform, built for good reasons and working.
Absent one of those, building is usually a preference expressed as a requirement. The test worth applying: would a customer ever notice that you built the gateway yourself? If not, it is infrastructure, and infrastructure is bought.
15. Twelve mistakes
- Building the platform before shipping a feature. Months of work with nothing to show.
- Choosing on the demonstration. All demonstrations are impressive; none discriminate.
- No evaluation set. No way to know whether anything improved.
- Prompts stored only in the platform. The most portable asset made unportable.
- Application code calling the platform directly. Every integration point becomes migration work.
- Defaulting to the most capable model. The largest avoidable cost.
- Cost visible only in their dashboard. Cannot attribute, cannot optimise.
- Budgets that alert rather than reject. A metric during the incident that matters.
- No degradation path. Their outage becomes your outage.
- Ignoring subprocessors. A compliance gap discovered during a review.
- Starting with the highest-stakes use case. Where an early failure is most expensive politically.
- Assuming buying transfers accountability. It does not, in any jurisdiction.
16. A worked example: two teams, same feature
Consider two teams at similar companies, both building the same thing: a feature that answers customer questions from their own documentation, drafted for an agent to review.
The building team starts with infrastructure. They spend the first two months on a gateway, prompt versioning, a retrieval pipeline and basic observability. It is competent work and none of it is visible to anyone outside engineering. In month three they ship a first version. In month four they discover their retrieval returns plausible-but-wrong passages and add reranking, which they had not known to include. In month six the feature is good, and they have built a platform they now maintain.
The buying team ships a first version in week three. Not because the platform is magic, but because the two months of plumbing did not happen. Reranking is on by default, so they never encounter the problem the other team spent a month diagnosing. What they spend their first two months on instead is the documentation itself — restructuring content that was written for humans browsing rather than for retrieval, adding the metadata that makes permission filtering possible, and fixing the fact that a third of the corpus was out of date.
Both teams arrive at a working feature at roughly the same time, which is the uncomfortable finding. But the buying team spent that time on their corpus and their evaluation set, and the building team spent it on infrastructure. Six months later the buying team's answer quality is materially better, because retrieval quality caps everything and they improved the thing that actually caps it.
The cost picture also inverts from what was expected. The building team pays provider rates directly and routes everything to the most capable model, because implementing routing was scheduled for later and later never came. The buying team pays a margin on inference and routes most requests to a cheaper model with caching enabled by default. Per resolved question, the buying team pays less despite the margin.
Where the building team is genuinely ahead is in a specific capability they needed — an unusual permission model that the platform expresses awkwardly. They built exactly what they required. Whether that advantage justified four months of platform work is the question, and for most organisations the honest answer is no.
The generalisable point: the platform work is not the hard part and it is not the differentiating part. The corpus, the evaluation set and the product decisions are both, and they take the same amount of time regardless of who built the gateway.
17. A capability checklist for the evaluation
Vendor comparisons drift toward whatever the vendor emphasises. A fixed checklist keeps the evaluation honest, and the questions below are the ones that separate platforms in practice rather than on a feature grid.
| Area | The question to ask | Why it matters |
|---|---|---|
| Model range | Do they offer a genuine range from cheap and fast to slow and capable? | Most workloads need both. A single-tier provider forces you to overpay on volume tasks or under-serve hard ones. |
| Model lifecycle | What notice do you get before a model is retired, and how long do old versions stay available? | A silently changing model breaks evaluations you thought were stable. Pinned versions with a published deprecation window are what you actually need. |
| Latency profile | What is the time to first token, not just total latency? | For anything streamed to a user, time to first token is what the experience feels like. Total time matters for batch work and almost nothing else. |
| Throughput limits | What are the rate limits, how do they scale, and what happens when you hit them? | A limit that returns a clear retriable error is manageable. One that silently queues for thirty seconds is an outage in disguise. |
| Cost controls | Can you cap spend per key, per project, per period? | Without caps, one bad loop is an unbounded bill. This is the control most teams wish they had configured before they needed it. |
| Caching | Is there a mechanism to avoid reprocessing a stable prefix? | For long system prompts or attached documents, this is frequently the largest single cost reduction available. |
| Batch processing | Is there a discounted path for work that tolerates delay? | Classification, enrichment and backfills rarely need to be immediate, and the discount is substantial. |
| Data handling | Is your data used for training, and how long is it retained? | The answer belongs in the contract, not the marketing page. "No training on customer data" and "zero retention" are different commitments. |
| Residency | Can you pin processing to a region? | Non-negotiable for some regulated workloads and irrelevant for others. Establish which you are early. |
| Certifications | Which attestations exist, and are the reports actually obtainable? | A logo on a page is not a report. Ask for the document before you commit. |
| Observability | Can you see per-request cost, latency and token counts, exported somewhere you control? | Without this you cannot attribute spend, and unattributed spend is spend nobody optimises. |
| Status transparency | Is there a real status page with history, and do incidents appear on it promptly? | A provider whose status page is always green during your outages is telling you something. |
| Interface stability | How compatible is the interface with alternatives, and how much of your code is provider-specific? | This determines what switching actually costs, and it is usually far less than teams fear. |
| Support | Who do you reach at three in the morning, and what have they committed to? | Self-serve support is fine until the feature is load-bearing. |
Two pieces of advice on running the evaluation itself. First, test on your own data, not on the vendor's examples — published benchmarks tell you very little about how a model handles your specific domain, your document formats and your edge cases. Build a set of fifty representative cases with known good answers and run every candidate against it. That afternoon of work is worth more than a month of reading comparisons.
Second, test the failure paths deliberately: exceed the rate limit and see what comes back, send a malformed request, send something that will be refused, and send something far larger than your typical input. How a platform behaves when things go wrong is a better predictor of how much it will cost you operationally than how it behaves when they go right.
18. Frequently asked questions
Is a platform layer worth the margin on inference?
Usually, and frequently it is cost-negative rather than cost-positive — routing and caching that a platform does by default typically save more than the margin costs. The margin becomes material at sustained high volume, which is the point at which building starts to make sense and also the point at which you know your requirements well enough to build the right thing.
How do we avoid lock-in?
Keep prompts and evaluation sets in your repository, store source documents in your own systems, define your own interface so application code never calls the platform directly, and record per-request detail in your own telemetry. Those four cost almost nothing and turn a migration into a project rather than a rewrite. Refusing to use platform features at all is not avoiding lock-in; it is declining the value you are paying for.
Can we send customer data to one of these?
Subject to your obligations and their documented terms — which is a procurement question with answers rather than an unknowable risk. Confirm retention, training use, processing location and the subprocessor chain in writing. Many organisations run regulated workloads on managed platforms under agreements that address exactly this; the ones that struggle are those that never asked.
What if the provider is acquired or shuts down?
The same risk as any vendor, mitigated the same way: keep your portable assets portable, know what migrating would involve, and prefer providers whose commercial position looks durable. This is a real consideration in a young market, and it is an argument for keeping the exit open rather than for building everything yourself.
Should we use one platform for everything?
Decide per use case rather than as policy. An applied product for a commodity capability, a platform layer for features that matter, and direct model access for anything unusual is a common and sensible arrangement. What causes trouble is running two full platforms in parallel without a rule for which is used when.
How long should a pilot run?
Long enough to build an evaluation set from real traffic — typically four to six weeks with genuine users. Shorter pilots measure the demonstration rather than the product. Whatever you decide afterwards, the evaluation set is the thing you keep, which makes the pilot worthwhile even if the answer is no.
What is the most common reason adoption disappoints?
Expecting the platform to solve the data problem. Retrieval quality caps answer quality, and retrieval quality is a function of how well your content is structured, permissioned and kept current — none of which a platform can do for you. Teams that budget for corpus work are satisfied; teams that expect to point a platform at a document folder are not.
What should we do first?
Settle the data policy, then run a real pilot on one narrow use case with a human in the loop, measuring against an outcome you named in advance. Build the evaluation set during it. That takes about six weeks and produces both a decision and the asset you need regardless of which way the decision goes.
Key takeaways
- The platform work is required either way. The only question is whether building it is a differentiator or a tax.
- Buy the machinery, own the corpus and the evaluation set. Those are the assets that determine quality and they are portable.
- The omitted cost is opportunity cost. Months spent on a gateway are months not spent on the product.
- Retrieval quality caps everything, and no platform fixes your content for you.
- Buying does not transfer accountability. A wrong answer to a customer is yours.
- Keep the exit open cheaply: your interface, your prompts, your documents, your telemetry.
The organisations shipping useful AI features fastest are rarely the ones with the most sophisticated internal platforms. They are the ones who rented the plumbing, spent the saved months on their data and their evaluation set, and put something in front of real users early enough to learn what it actually needed to do.
Enjoyed this article?
Get more engineering insights from ELIVTECH — or talk to us about your project.
Get in touch