Skip to content
Back to Blog
Linux & Server10 min read

Cloud Infrastructure Setup: 5 Decisions Before You Deploy

Region selection, network topology, IAM policies, backup automation, and monitoring infrastructure—these five architectural choices define stability, cost, and recovery before your first resource goes live.

Written by Abdul AbrorTechnical Hosting Support Engineer
Cloud Infrastructure Setup: 5 Decisions Before You Deploy
On this page

Most infrastructure problems I've debugged trace back to day-zero architecture choices. A subnet range that looked generous in month one becomes a hard ceiling in month six. An IAM policy that seemed "simple" metastasizes into a maintenance nightmare.

You can refactor later, sure. But migrations always cost more than planning.

This article walks through the five architectural decisions that set the boundaries for every cloud deployment: where you run, how traffic flows, who can do what, how you recover from failure, and how you see what's happening. Get these right and you've built a foundation that scales. Get them wrong and you're rewriting Terraform at 2 a.m.

Decision 1: Region and Availability Zone Strategy

Pick your region first. Latency to end users matters, but so does cost.

Pricing varies significantly between regions—sometimes 20-30% for identical instance types. Balance that against round-trip time. If your users are in Europe and you deploy in Asia-Pacific to save eight dollars a month, you've traded pennies for hundreds of milliseconds.

Compliance often decides for you. GDPR workloads need EU regions. Healthcare data under HIPAA has residency requirements. Financial services face similar constraints. Check your regulatory map before you check the pricing calculator.

Availability zones are the next layer. Most providers offer three or more AZs per region, each with independent power and networking.

Deploy across at least two AZs for anything production. Single-AZ deployments will fail, and they'll fail at the worst possible time—during a provider-side outage when support queues are measured in hours.

Multi-region or single-region?

Multi-region setups buy you disaster recovery and geo-distribution. They also double your operational complexity.

In support tickets I handled, teams that deployed multi-region without a clear DR plan or traffic-management strategy ended up with two half-working environments instead of one reliable one. Data consistency across regions is hard. DNS failover sounds simple until you test it.

Start single-region unless you have a concrete reason—measured latency requirements, regulatory data residency in multiple jurisdictions, or an RTO that requires instant geographic failover. You can always expand later.

Decision 2: Network Topology and Segmentation

Your network design controls blast radius. A flat network where every resource can reach every other resource is a security auditor's nightmare and an attacker's playground.

Subnet strategy

Plan your CIDR blocks with room to grow. I've seen teams start with a /24 (254 hosts) thinking it's "plenty," then hit the ceiling when autoscaling spins up during a traffic spike.

Use at least a /16 for your VPC. Carve out smaller subnets for different tiers:

10.0.0.0/16 - VPC range
  10.0.1.0/24 - Public subnet AZ-A
  10.0.2.0/24 - Public subnet AZ-B
  10.0.10.0/24 - Private app subnet AZ-A
  10.0.11.0/24 - Private app subnet AZ-B
  10.0.20.0/24 - Private database subnet AZ-A
  10.0.21.0/24 - Private database subnet AZ-B

Public subnets hold load balancers and NAT gateways. Application servers live in private subnets with no direct internet access. Databases sit in their own isolated tier.

This three-tier model is standard because it works. Public-facing components are minimized, application logic runs behind a network boundary, and your data layer has two levels of isolation.

Routing and gateways

NAT gateways let private instances reach the internet for updates without exposing inbound ports. They're billed by the hour and by data processed, so they're not free. But they're cheaper than a breach.

Internet gateways attach to your VPC and route public subnet traffic. One IGW per VPC is typical. Keep it simple.

Transit gateways or VPC peering connect multiple VPCs. Peering is simpler for a handful of VPCs. Transit gateways scale better but add cost and a central failure point. Don't overengineer this until you actually have multiple VPCs that need to talk.

Security groups vs. network ACLs

Security groups are stateful firewalls at the instance level. They're where most of your rules live. Start with deny-all and open only what you need:

# Web tier security group
Ingress: 443 from 0.0.0.0/0
Ingress: 80 from 0.0.0.0/0
Ingress: 22 from bastion-sg
Egress: all to app-tier-sg

# App tier security group  
Ingress: 8080 from web-tier-sg
Egress: 5432 to db-tier-sg
Egress: 443 to 0.0.0.0/0 (for API calls)

# Database tier security group
Ingress: 5432 from app-tier-sg
Egress: none

Network ACLs are stateless and operate at the subnet boundary. Most teams set them once and never touch them. Use security groups for day-to-day rules.

Decision 3: IAM Model and Permission Boundaries

Identity and access management is where most cloud accounts leak permissions. The principle is simple: grant minimum necessary access. The execution is harder.

Human users vs. service accounts

Humans log in with SSO or federated identity. Service accounts (roles, service principals, workload identities) are attached to compute resources.

Never put long-lived credentials in code or config files. Use instance roles or workload identity instead. If your application needs to read from object storage, attach a role with read-only permissions to that bucket. The SDK picks up credentials automatically.

For AWS, that looks like:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3:GetObject"],
    "Resource": "arn:aws:s3:::my-app-bucket/*"
  }]
}

Attach that policy to an IAM role, assign the role to your EC2 instance or Lambda function, and you're done. No access keys in environment variables.

Organize with groups and roles

Create groups for common access patterns: developers, operators, read-only auditors. Assign users to groups, not individual permissions.

Service roles follow the same logic. If ten Lambda functions all need the same DynamoDB permissions, create one role and share it.

Permission boundaries and SCPs

Permission boundaries set a maximum ceiling. Even if a user has admin rights within their boundary, they can't escalate beyond it. Useful for delegating limited admin access to teams.

Service control policies (in AWS Organizations) enforce rules across accounts. You can block entire regions, deny root user access, or require MFA. SCPs trump everything else.

In a typical setup, you'd use an SCP to block all regions except the two you've chosen, deny deletion of CloudTrail logs, and require encryption for object storage. That's your guardrail. Teams operate inside it.

Decision 4: Backup and Disaster Recovery Architecture

Backups are boring until you need them. DR is expensive until you calculate downtime cost.

RTO and RPO

Recovery time objective: how long can you be down? Recovery point objective: how much data can you lose?

If your RTO is four hours and your RPO is 15 minutes, you need frequent backups and a restore process you've tested to complete in under four hours. That's your design constraint.

Most teams overestimate their RTO. "We can be down for a day" sounds fine until the CEO is in your inbox at hour two.

Snapshot strategy

Block storage snapshots are incremental and cheap. Schedule them daily for most workloads, hourly for databases or high-change volumes.

Retention depends on RPO. Keep seven daily snapshots, four weeklies, and 12 monthlies for a typical retention curve. Adjust based on compliance requirements.

Tag your snapshots with metadata—instance ID, environment, date. You'll thank yourself when you're searching for the right one during an incident.

Database backups

Managed database services handle this for you. Enable automated backups, set your retention window, and configure your backup window for off-peak hours.

For self-managed databases, use logical dumps plus point-in-time recovery if supported. PostgreSQL gets pg_dump daily and WAL archiving to object storage. MySQL gets mysqldump and binary log shipping.

Test your restores. I've seen teams with perfect backup schedules discover during an outage that their restore process is broken or takes six times longer than expected.

Cross-region replication

For true DR, replicate backups to a second region. Object storage makes this straightforward—enable cross-region replication and your snapshots copy automatically.

Database replicas are harder. Managed services offer read replicas in other regions with a few clicks. Self-managed setups require streaming replication config and careful monitoring of replication lag.

Decision 5: Observability and Monitoring Setup

You can't fix what you can't see. So what should you watch?

Metrics, logs, traces

Metrics are time-series data: CPU, memory, request rate, error count. Logs are event records. Traces follow a single request through multiple services.

For a typical web application: - Collect system metrics (CPU, RAM, disk, network) from every instance
- Collect application metrics (request latency, error rate, queue depth)
- Ship logs to a central store (CloudWatch, Elasticsearch, Loki)
- Instrument critical paths with distributed tracing

What to alert on

Alert on symptoms, not causes. "API error rate above 5%" is actionable. "Disk I/O wait above 10%" might matter or might not.

Start with these: - HTTP 5xx rate crosses threshold
- Request latency p99 exceeds SLA
- Any service becomes unreachable
- Disk usage above 80%
- SSL certificate expires in under 7 days

Avoid alert fatigue. If your team ignores 90% of alerts, you've trained them to ignore the one that matters.

Log aggregation

Centralize logs from all instances. Searching logs across 50 servers individually is not a plan.

Structure your logs. JSON is parsable. Plain text with inconsistent formatting is not.

{
  "timestamp": "2026-09-27T07:05:01Z",
  "level": "error",
  "service": "api-gateway",
  "message": "database connection timeout",
  "user_id": "12345",
  "request_id": "abc-def-ghi"
}

Now you can query by service, filter by error level, and trace a single request across services using the request ID.

Dashboards and runbooks

Build dashboards for each service showing key metrics. Your on-call engineer should be able to glance at a screen and know if the system is healthy.

Write runbooks for common alerts. "API error rate high" should link to a doc that says: check database connections, check external API status, review recent deploys, examine error logs for patterns.

Runbooks turn 3 a.m. panic into a checklist.

What To Lock Down First

Once you've made these five decisions, your first deploy should include: MFA enabled on all accounts, CloudTrail or equivalent audit logging turned on, encryption at rest for all storage, security groups locked to minimum necessary ports, and automated snapshots running.

Those five controls prevent the majority of avoidable incidents. Everything else is iteration.

Architecture decisions are constraints you choose. Choose well and they guide you toward stability. Choose poorly and they become technical debt with a monthly bill attached.

FAQ

How many environments should I plan for?

Three: production, staging, and development. Staging mirrors production config for final testing. Development is cheaper and more permissive for daily work. Some teams add a fourth for QA or performance testing.

Should I use managed services or run my own?

Managed services cost more per unit but save operational time. For databases, caches, and message queues, the tradeoff almost always favors managed unless you have specific performance or compliance needs. For compute, it depends on scale and control requirements.

When do I need a VPN or bastion host?

If you're SSHing into instances or accessing internal services from outside the VPC, you need secure access. Bastion hosts sit in a public subnet and act as a jump box. VPNs provide network-level access. Both work; bastions are simpler for small teams.

How do I estimate costs before deploying?

Use the provider's pricing calculator with realistic usage numbers. Compute, storage, and data transfer are your big three. Monitor spend daily in the first month and set budget alerts.

Can I change my network design later?

Yes, but it's painful. Changing CIDR blocks usually means creating a new VPC and migrating resources. Updating security group rules or adding subnets is easier. Front-load the planning.