Three Environments, One Oracle VM: Running Production, Staging and Test on a Single Box
SwitchWithAI's backend runs on one Oracle Cloud Ampere VM. Not one per environment: one. Production, staging and a test sandbox share the same four cores and 24 GB of memory.
That sounds like a shortcut, and in some ways it is. But a small product rarely needs three machines. What it does need is three environments that can never be confused with each other. This post is about how the setup keeps them apart, and about the one failure that made me take that seriously.
The three environments
| Environment | Role | Port |
|---|---|---|
| Production | Paying users, managed Postgres, tailoring only | 8081 |
| Staging | A mirror of production with its own throwaway database. The only route to production | 8082 |
| Test | A sandbox for experiments: the older job-hunt pipeline, agents, a headless browser | 8080 |
Each one has its own git checkout, its own environment file, its own Docker Compose project (so its own volumes and database), its own loopback port and its own nginx server block. They share hardware and nothing else.
The failure that shapes everything: the wrong site answers with a 200
Early on, ports were defaults. Two stacks could, in principle, race for the same one.
A port collision does not look like an outage. It looks like a healthy site. The request arrives, something answers, the status is 200. It is simply the wrong backend, talking to the wrong database. You can click around staging for ten minutes before noticing the data is production's, or the other way round.
That is why ports are now pinned per environment in a single file, deploy/envs.sh, which every deploy script sources. It is a plain lookup table that is not allowed to touch the network or the box, so it is always safe to read. The numbers are the ones the machine actually used when I wrote them down, not a tidier renumbering. Moving a running stack's port buys nothing and risks an outage.
One source of truth for "what does staging mean"
Before envs.sh, the meaning of "staging" was spread across a bootstrap script, an update script, an nginx template and my memory. Each copy was right when it was written.
Now bootstrap, update and status all read the same table: domain, checkout path, compose project, port, nginx template, branch, whether search engines may index it. swa.sh status prints all three environments side by side. When something looks off, that is the first command.
Production's edge is not the box's nginx
This surprised me the first time I traced it properly:
- The frontend is not on the VM at all. It is on Vercel.
- The API reaches the VM through a Cloudflare Tunnel.
cloudflaredon the box forwards to an nginx listener bound to loopback, which forwards to the production container.
The nginx hop exists for one job: blocking a set of legacy routes from an older product so they return 404 at the edge, whatever the application thinks. It passes the forwarded-for header through untouched, so the application's count of trusted proxy hops did not change.
Staging and test are served by the same nginx directly, with Let's Encrypt certificates. The old production nginx template still exists, but only as a rollback path.
Staging is the only way to production
Nothing is deployed to production by hand from a laptop. A change lands on staging, gets exercised against staging's own database, and is promoted from an admin Release tab. That single path is slower than git pull on the server, and it is the reason I sleep through deploys.
Living within 24 GB
Sharing a machine means every new thing has to declare its memory ceiling. The test environment runs a headless browser, which needs more shared memory than Docker's default or it crashes, and gets an explicit limit. The warm LibreOffice pool used for PDFs is capped at one or two workers, each a few hundred megabytes. Anything added to the box gets a cap before it gets a feature.
Would I do it again?
Yes, with one condition: the separation has to be enforced by the files, not by discipline. Pinned ports, one environment table, independent compose projects, and a single promotion path. With those, one machine behaves like three. Without them, it behaves like one machine that occasionally lies to you with a 200.
The product this runs is SwitchWithAI, an AI resume tailor that keeps your original formatting.