This site only uses technical cookies required for it to work: no tracking, no profiling. Cookie Policy

Skip to content
All terms

Serverless, on-prem or edge: where should workloads run?

Three deployment models: pay per execution in the cloud, own the infrastructure in house, or bring compute close to the data.

These are the three deployment models to choose from today, and the choice became interesting again precisely because of AI. Serverless: code runs in the cloud only when needed and you pay per execution, with no servers to manage. On-prem: the infrastructure is yours, on your premises or in a data center you control. Edge: compute sits close to where data is born, on the factory floor, in the store, on the device. The choice among the three is never purely technical: behind it sits a question about how predictable a workload is, how much response speed matters, where data can legally reside and what operational skills a team genuinely has in house. It is precisely the wave of AI workloads, often constant and high volume, that has made relevant again a comparison many companies had considered settled, after years in which the public cloud looked like the only sensible answer.

The criteria that decide

Four axes. The load profile: serverless excels at intermittent, unpredictable loads (you pay zero when nothing runs), while a constant heavy load, typical of high-volume AI inference, can cost more on serverless than dedicated infrastructure: one of the springs of cloud repatriation. Latency: if the answer must arrive in milliseconds next to the production line, edge has no rivals. Data: residency, sovereignty and confidentiality requirements push on-prem or edge (data never leaves the perimeter), a theme the CLOUD Act made concrete. People: on-prem requires operational skills that serverless eliminates; underestimating that is the most expensive mistake in price-list-only comparisons.

The honest answer

Almost no serious company is 100% on any of the three: the norm is hybrid, decided workload by workload. A recurring pattern in enterprise AI: prototypes and spikes on serverless, steady inference on dedicated infrastructure (cloud or on-prem, with FinOps numbers in hand), and small models at the edge where latency or confidentiality demand it. The right question is not "which model wins" but "this workload, with this load profile, this data and this team, where does it cost less and risk less?".

  • Cloud repatriation · The selective move of workloads from public cloud back to on-premise or hybrid environments, for cost and control. A FinOps decision, not a retreat.
  • FinOps · The practice bringing financial accountability to the cloud: every team sees, understands and optimizes the cost of what it runs.
  • Inference · Using an already trained AI model to produce answers: every ChatGPT question is inference, and it is where costs concentrate today.
  • CLOUD Act · US law compelling American providers to hand over data under US legal orders, even when it is stored in Europe.

A term that hits close to home? Let's talk.

CONTACT ME