Cloud Architecture

Cloud architecture that grows with your product, not ahead of it.

Decisions made in the first year of a product, such as how services are split, where data lives and how the system scales, tend to stay for many years. We help growing companies design on AWS, Google Cloud or Azure with that in mind, including GPU and inference workloads, vector search and the data-handling questions AI features raise.

What we design

The decisions a cloud architecture consultant should help you get right

We work outward from your product, team and growth plans. The goal is the simplest architecture that will comfortably carry the next stage of the business.

Service boundaries

Whether a well-structured monolith or separate services suits your team size, and where the natural seams are if you split later.

Data stores and data flow

Relational, document, cache, search and vector stores chosen for real access patterns, with a clear owner for each piece of data.

Scaling and failure modes

What happens at ten times today's load, when a zone or third-party API fails, and which parts can degrade gracefully.

Security boundaries

Network isolation, identity and access policies, encryption and least-privilege roles designed in, not bolted on the week before an audit.

Regions and data residency

Where customer data is stored and processed, which matters for regulated industries and for AI providers handling personal data.

AI workloads

Room for inference endpoints, GPU capacity, embedding pipelines and model API traffic without rebuilding the platform.

Designing for AI

AI workloads change the shape of a cloud architecture.

A conventional web product has fairly predictable compute and storage needs. AI features bring different characteristics: GPU instances that are expensive and sometimes scarce, inference calls with long and variable response times, embedding pipelines that churn through large volumes of documents, and vector search sitting alongside your main database. Each one affects how you design queues, timeouts, caching and capacity.

There are also choices that did not exist a few years ago. Call a hosted model API, run open models on managed GPU services, or host them in your own environment for data control? Keep embeddings in PostgreSQL with a vector extension, or add a dedicated vector store? The right answers depend on data sensitivity, volume and latency, and they are far easier to revisit when the architecture keeps AI behind clear interfaces.

We design for the AI workloads you have and leave sensible room for the ones on your roadmap, without paying for GPU capacity you do not need yet. Once a system is running, trimming an existing bill is its own discipline, covered by our cloud cost optimisation work.

Compute options

Serverless, managed containers or Kubernetes?

The compute model is one of the first architecture decisions and one of the most argued about. For growing products, the simplest option that meets your needs usually wins.

Serverless functionsManaged containersKubernetes
Operational effortLowestLow to moderateHighest, needs dedicated skills
Cost at low trafficVery low, pay per requestModest baselineHigher baseline for the cluster
Cost at steady high trafficCan become expensivePredictableEfficient when run well
Long-running and AI jobsConstrained by timeouts and memoryGood fit, with GPU options on some platformsMost flexible for GPU scheduling
ExamplesAWS Lambda, Cloud Run functions, Azure FunctionsAmazon ECS, Cloud Run, Azure Container AppsEKS, GKE, AKS
Typical fitEvent-driven tasks and spiky APIsMost product backendsMany services and a platform team

We have no ties to any cloud provider. AWS, Google Cloud and Azure are all capable choices, and your team's existing experience often matters more than a feature list.

How we work

From review to an architecture your team can build

A cloud architecture engagement usually runs as an advisory sprint of 2 to 4 weeks, and can continue into hands-on implementation with your team or ours.

Understand the product

Week 1

Growth expectations, compliance needs, AI plans, team skills and the current setup, including what already works and should be left alone.

Design the options

Weeks 1 to 2

Two or three candidate architectures with trade-offs in cost, complexity and risk. AI helps us model scenarios and draft diagrams quickly, while senior engineers make the calls.

Prove the risky parts

Weeks 2 to 3

Small proofs of concept for the uncertain pieces, such as a GPU inference path, a data migration or behaviour under load, before anyone commits.

Document and hand over

Weeks 3 to 4

Architecture decision records, infrastructure-as-code foundations and a phased plan your engineers can build from.

FAQ

Questions we often hear

What does a cloud architecture consultant do?

A cloud architecture consultant designs how your system runs in the cloud: compute, data stores, networking, security, scaling and resilience. They balance cost, complexity and growth plans, and document decisions so your team can build and operate the result. Good ones also tell you what not to build yet.

Should we use AWS, Google Cloud or Azure?

All three are capable for most products. The deciding factors are usually your team's existing experience, specific managed services you rely on, credits or commercial terms, where your customers are and, for AI workloads, which models and GPU options you can access. Switching later is costly, so it deserves a considered decision rather than a default.

How do you design cloud architecture for AI workloads?

Put AI behind clear interfaces so you can move between hosted model APIs, managed GPU services and self-hosted models. Plan for long, variable inference times with queues and timeouts, decide where documents and embeddings are stored, and confirm which data may leave your environment. Size GPU capacity to real usage rather than projections.

Do we need Kubernetes?

Most growing products do not. Managed container services or serverless platforms meet the needs of many teams with far less operational work. Kubernetes starts to make sense with many services, specialised scheduling such as GPU workloads, or a team with the skills to run it.

Is cloud architecture consulting the same as cloud cost optimisation?

No. Architecture consulting is about designing systems well, which shapes cost among many other things. Cloud cost optimisation focuses on reducing the bill of a system that already runs, through rightsizing, commitments and removing waste. If your main problem is an existing bill, start there.

Design with confidence

What does your architecture need to support?

Share what your product does, how it is hosted today, and your growth or AI plans. We will reply within 24 hours with the architecture questions we would tackle first, and an honest view on whether you need outside help at all.

Working with companies globally · Response within 24 hours