What counts as activity
Only real traffic:- an HTTP request, stamped when it arrives and again every 5 seconds while it is still open, so a long download or a websocket stays active for its whole life
- any byte read on a database lane, in either direction
- coming up after a deploy or a wake
metrics and logs polling (insta compute metrics,
insta compute logs, and their postgres/redis/mysql/mongodb equivalents) do not count, and the
dashboard polling its own pages does not count. If it did, an open browser tab would keep an entire
box awake.
When a service sleeps
The sweeper looks at every service every 30 seconds and sleeps one when all of these hold:- the container is running
- the service is not always-on
- its desired state is running (something you stopped by hand stays stopped)
- the last activity is at least the idle window ago, which is 5 minutes for compute and 10 minutes for databases
- the row was created at least the create grace ago, 10 minutes by default
- no deploy, wake or other operation is in flight for it
Memory pressure
Sleeping on idle is not always fast enough, so the sweeper also watches free memory. While available memory is below 15 percent of the total, it sleeps the least recently active running service, one at a time, until it is back above the floor. Never evicted: an always-on service, a service that woke inside the wake-protection window (60 seconds by default), and a service answering a request right now. A wake asks for room BEFORE it starts anything. It works out what the waking service needs, sleeps the least recently active candidates until that much is free above the floor, and starts the container only then, so the victim’s memory is released before the new one claims any. When there is already room this costs nothing, because the check returns before it reads or stops anything. When there is not, the wake pays one stop grace PER SERVICE it has to sleep, because it evicts in a loop until there is room, and that is the moment that price is worth paying: a slow wake beats a box in swap-free reclaim. Two outcomes here are not the same, and it matters which one you are looking at:- Nothing left to evict. Every other service is always-on, answering a request right now, or inside its wake-protection window. The scheduler logs that once and the wake PROCEEDS, deliberately: refusing would make a service permanently unreachable on a box that is merely full, and the kernel is the better last resort than a daemon that will not start anything. This is what you will see on a small box whose services genuinely do not fit together.
- Eviction itself failed. A stop that errors, not an empty pool. That is a fault in making
room rather than an absence of room, so the wake fails with
could not make room to wake <container>instead of starting into a floor it could not clear.
/proc/meminfo, or the daemon’s own cgroup ceiling
when it runs in a container that has one, whichever is smaller. A ceiling matters because
/proc/meminfo describes the machine even from inside a container, so a memory-limited daemon
would otherwise believe it had the whole box and evict nothing.
On ZFS the reading adds back the reclaimable part of the ARC. The ARC is not page cache, so the
kernel’s MemAvailable leaves it out even though it is handed back under pressure, and a host
with a warm ARC would otherwise read as permanently under the floor and evict on every sweep. See
ZFS.
On macOS the reading Docker reports is not the host memory, so pressure eviction is off unless you
set INSTA_OSS_MEM_BUDGET_MB to the budget you want the daemon to assume.
What sleep does to your process
Sleep isdocker stop, which sends the stop signal the image declares: SIGTERM for apps and for
managed databases, SIGINT for the official Postgres image, which is its fast shutdown. Compute
containers run with --init, so the signal reaches your process even when the entrypoint is a
shell script.
The daemon then waits 10 seconds, or 30 for a database, before SIGKILL. Handle the signal if you
have work to flush.
What wakes it
- A request on any lane. HTTP, Postgres, Redis, MongoDB or MySQL. The connection is held while the container starts, for up to 60 seconds, plus up to 30 seconds of readiness retries after that. You see a slow first request, not an error. Readiness here means the port accepts a connection. That is the whole test for an HTTP service, and it is not the same thing as the app being ready to answer: an app that binds its port early and mounts its routes later will take the forwarded request and answer it out of its own half-built stack, so the first page after a wake can come back 404 or 502 from the app rather than slow. The declared template healthcheck gates a deploy, not a wake, and for a slow app it would not close the window anyway: the measured case answered its own healthcheck several seconds before it would serve a page. Always-on is the answer, which is why the bundled templates that boot slowly declare it. Databases are probed on the wire instead, so they do not have this window.
-
insta compute start, and any management call that needs the service up. - A deploy, which starts the replacement container.
insta compute stop is different: it is a durable intent, and traffic never wakes a service you
stopped. Requests answer 503 until you run insta compute start.
Always-on
For anything that has to keep running with no inbound traffic, a queue worker or a scheduler:The defaults match the hosted platform: compute on the default branch is always-on and
scale-to-zero is opt-in, managed Redis, MySQL and MongoDB follow the same default, and Postgres
scales to zero. The one addition for a single box is that branch clones scale to zero unless a
service is explicitly always-on, so forking a project does not start every app it has. Switch a
service off with
insta compute always-on off web; send {"enabled": null} to
PUT /projects/<id>/services/<sid>/always-on to put it back on the default; or set
INSTA_OSS_ALWAYS_ON_DEFAULT=0 to make default-branch services scale to zero too. A service
that becomes always-on while it is asleep, and was not stopped by hand, is woken by the next sweep
(every 30 seconds by default, INSTA_OSS_SWEEP_SEC); with the sweeper off
(INSTA_OSS_SCHEDULER=0) it stays asleep until its next request.Tuning
Set these in/etc/instacloud/instad.env, then cd /etc/instacloud && docker compose up -d.
An idle window of 0 turns sleeping off for that kind of service without turning off the sweeper.
How it reads
Why a public app keeps waking up
On a box with a public domain this is the first thing that looks like a bug, and it is not one. Unless you supplied a wildcard (--tls custom, in the fixes below), the edge issues a
certificate the first time somebody asks for a hostname over HTTPS. Every
certificate a public authority issues is published to the certificate transparency logs, and those
logs are read continuously by scanners. Within minutes of that first request, more requests start
arriving from hosts nobody invited. They are ordinary HTTP requests, so they count as activity, so
they wake the service, and an operator watching docker ps sees containers coming back on their
own and concludes scale to zero is broken.
We measured it on a public box, and the numbers say more than the mechanism does. A freshly
issued service hostname drew 141 requests from one credential scanner probing /.env,
/.env.local, /config.js and /settings.js within about 15 minutes, and the requests kept
arriving every 1 to 3 minutes, with the largest gap being 3 minutes. The compute idle window is
300 seconds and any request resets it. So this is not an occasional extra wake: on a
per-hostname certificate, a public compute service never sleeps at all. A control service whose
hostname never received a public certificate slept on schedule twice in a row. Databases keep
sleeping either way, because their lanes are not discoverable from a transparency log.
The fix is the certificate, not the timer, and it is the same thing the managed cloud does: serve
ONE wildcard certificate for *.<domain> so that deploying a service publishes nothing.
- Supply a wildcard.
--tls customwith a certificate for*.<domain>and*.s3.<domain>(buckets are addressed<bucket>.s3.<domain>, and a wildcard covers one label) issues nothing per hostname, so no name is published, nothing arrives to find it, and the idle timer runs to the end. See TLS and what your domain costs you, which also covers what you take on: renewal is yours. - Or keep the box off public certificates.
--tls internaluses Caddy’s own CA, which publishes nothing; no browser trusts it without installing the root, which is fine for a box you reach yourself. - The auto sslip.io domain is the case that bites. It exists to let you try the product in one command, and it puts every hostname you deploy into a public log. Do not size a box, or judge scale-to-zero, on that configuration.
- To see what woke something, read
insta agent events --json. Everyservice.wakecarries the door it came in on,traffic,apiordeploy, so a run oftrafficwakes with no deploy near them is the scanners. The event does not record who sent the request;insta compute logson the service does.