Second transit carrier live in Chișinău — 20 Gbps of blended capacity. 20 Gbps blended uplink now live Why Moldova

Operations Practical

Moving a live site to a host you cannot ask for help

Nothing you install makes a migration seamless. What makes it bounded is an order of operations, and three of its steps happen days before you copy a single byte. Here is the sequence, the clock you do not control, and the honest arithmetic of the one window you cannot avoid.

18 min read Published 16 September 2026 Checked 23 days ago

“Migrate with zero downtime” is sold as a property of a tool, and it is not one. It is a property of an order of operations. No script, control panel or managed transfer service changes the fact that a domain points at exactly one address at a time, that the internet learned that address from you at some point in the past, and that it will keep using what it learned until the number you attached to it runs out. Almost everything that goes wrong in a migration is a step done in the right way at the wrong moment. So this is the sequence rather than the commands — including the three steps that happen days before you copy a byte, the one clock you are not allowed to touch, and the honest arithmetic of the window you cannot design away.

The clock starts before you copy anything

Begin with the thing that decides your whole schedule, because it is the one thing you cannot accelerate once you are standing in it. The domain name system does not push. There is no broadcast, no notification, no update that travels outwards from your zone. The specifications that define it — RFC 1034 and RFC 1035, unchanged in this respect since 1987 — describe a system that is asked and answers. Nothing tells the world your address changed. The world finds out by asking again, and it only asks again when the copy it is holding expires.

Which gives the single most useful sentence in this article: the time-to-live that governs your cutover is the one that was already published when a resolver last asked — not the one you are writing today. If your A record has been sitting at 86 400 seconds and you lower it to 300 an hour before the move, every resolver that fetched it in the past day is still holding a twenty-four-hour copy and is under no obligation to ask you anything for the rest of it. You did not shorten the wait. You shortened the wait for the people who happened to ask after you made the change, which is not the group you were worried about.

So the first action of a migration happens nowhere near the new server: lower the TTL, then wait at least one full old TTL before you do anything else. A day of patience buys you a five-minute cutover. Skipping it buys you a day of split traffic in which half your users are writing to a database you are about to abandon.

The line to remember: “DNS propagation” is not a process you wait out. It is a number you published in advance, and it is the only part of the schedule that is entirely your own doing.

Which means the question is never “how long does propagation take”. It is: what TTL was live on this record a day ago, and have I already served it long enough for the short one to replace it everywhere?

Now the part almost nobody accounts for. Your record’s TTL is yours. The TTL on the delegation above it is not. The NS records that tell the world which nameservers are authoritative for your domain are served by the registry of your top-level domain, with a TTL the registry chooses; for .com and .net that value is 172 800 seconds — two days — and no setting in your registrar panel lowers it. You cannot shorten it, you cannot flush it, and it applies from the moment you change nameservers.

The consequence is a rule with no exceptions: do not change DNS operator and server address in the same operation. If you do, you have deliberately created a two-day period in which some resolvers ask your old nameservers and some ask your new ones, and the only way to survive it is to keep both sets answering identically — which means you now have two zones to keep in sync during the exact week you are least able to notice they have drifted. Move the zone first, let the delegation settle for two days, confirm both operators agree, and only then start the actual migration. Two operations, two clocks, neither one interesting.

That ordering has a second consequence, and it is the one that catches people migrating away from a provider they have stopped trusting. If your current host also runs your DNS — or worse, holds the domain registration — then the zone move is not a nicety you can postpone until after the server move. It is the first migration, and it has to complete while your relationship with that provider is still cordial. Who holds which layer, and why they should not be the same company, is its own guide; the part that matters here is the sequence.

Three clocks in a migration, who controls each, and when it starts
What is cachedYours to setTypical durationWhen the clock actually starts, and what it costs to get wrong
Your A / AAAA record Yes 300 s once lowered, 3 600–86 400 s by default At the last fetch, under the old value. This is why the lowering has to precede the move by one full old TTL rather than by an hour. Cost of getting it wrong: a long tail of visitors served by a machine you have stopped updating.
The NS delegation above you No 172 800 s for .com and .net At the moment you change nameservers at the registrar. Nothing shortens it. Cost of getting it wrong: two days of two authoritative answers, which is survivable only if you planned for it.
A name that did not exist yet Partly The SOA minimum field At the first query that arrived too early. RFC 2308 has negative answers cached under the SOA minimum, not under the record’s TTL — so a hostname somebody looked up before you created it stays missing for that long. Cost: a subdomain that is “broken” for hours after you correctly created it.
The resolver in front of the resolver No Seconds to minutes Operating systems and browsers keep their own short caches, and a minority of access networks clamp very low TTLs upward to save queries. This is the reason a cutover is never instant even when you did everything right, and the reason you test with a hosts file rather than by asking a colleague to reload.

What has to move, and what quietly does not

Every migration has two inventories. The first is the one you write down without prompting: the files, the database, the web server configuration, the certificate. That pile is large, obvious, and almost never the thing that breaks — precisely because it is obvious. The second inventory is made of state that does not live in any directory you are about to copy, and it is where the outage comes from. It is worth writing this list out before you start, because nothing in it announces itself.

  • Scheduled work. Cron entries and systemd timers live outside your application directory, and a job that silently stops running is the failure mode nobody notices for a fortnight. Worse: if you copy them before cutover, both machines run them, and anything that sends mail or charges a card now does it twice.
  • Database users and grants. A dump contains your data. It does not contain the accounts that are allowed to read it, or the host each account is allowed to connect from.
  • Secrets that were never in the repository. Environment files, API keys pasted into a service unit, the credential somebody set by hand at three in the morning in 2023.
  • Reverse DNS. The PTR for a new address belongs to whoever holds the address, which is your new provider, and it is set in their panel rather than in your zone.
  • Firewall rules and fail2ban state. Rules built up over years, often as responses to specific incidents, and rarely written down anywhere else.
  • Third parties holding your old address. Payment webhooks, OAuth redirect allowlists, an upstream API that permits your IP and no other, a partner’s firewall. These fail after a successful cutover, which makes them hard to attribute, and several of them can only be changed by somebody who is not you and does not work weekends.
  • SSH host keys. The new machine has new ones. Your monitoring, your deployment pipeline and your own terminal will all object, and the objection looks exactly like the thing you should never click through.

Do the second inventory first, and do it as a written list rather than from memory. The first inventory copies itself with one command; the second one is the entire job.

Running both servers at once, on purpose

The instinct is to treat the overlap between two servers as waste. It is the opposite: the overlap is the whole safety mechanism, and one extra month of a server you are leaving is the cheapest insurance in this business. Provision the new machine, build it completely, and run it in parallel long enough to be bored by it. A migration you can abandon at any point up to the final minute is a different kind of event from one you have committed to.

Copy the bulk while everything is still live. For file trees, run the transfer once against a running system — it will take as long as it takes, and it does not matter, because nothing depends on it yet — then run it again at cutover, when it only has to carry what changed since. The first pass is hours; the second is minutes. That difference is the entire reason the window is short, and it costs nothing but a rehearsal.

Test the new server before the world sees it, and test it as the domain, not as an IP address. Point your own machine at the new address with a hosts-file entry and browse the real site under its real name. Anything that redirects, signs cookies, builds absolute URLs or checks its own hostname behaves differently under a bare IP, which means testing by IP proves the machine boots and nothing else. This is also how you avoid the tempting mistake of moving DNS “just for a minute to check” — a minute is not a unit the cache understands.

  1. The certificate has to exist before the cutover, and the usual method cannot produce it. The ACME HTTP-01 challenge proves control by fetching a file over the name being certified — which resolves to the old server, because that is the situation you are still in. It will fail, and it will fail in a way that reads like a firewall problem. Use the DNS-01 challenge, which proves control through a record in the zone and does not care where the name points, or copy the existing certificate and private key across and let the new machine renew normally afterwards. Do not loop a failing client: issuance is rate-limited per name and a retry storm spends a quota you will want at cutover.
  2. Decide what happens to the database, in writing. Files can be copied ahead of time because yesterday’s copy of a file is still a valid file. A database is a moving target, and there are exactly two honest answers: accept a short interval in which the site takes no writes, or set up replication from old to new and promote the replica. Replication is the right answer above a certain size and the wrong answer below it, because it adds a system you have never operated to the week you can least afford one.
  3. Disable the scheduled jobs on the new machine until it is live. Build them, then leave them switched off. Two servers running the same billing job against the same database is a worse outcome than an hour of downtime, and it is not detected by any check you are likely to have.
  4. Write the rollback before you need it. It is one line — put the old address back — and it works for exactly as long as the old server still has the data. The moment you accept writes on the new machine, rolling back means losing them. Know where that moment is, because it is the real point of no return, and it is not the same instant as the DNS change.

The cutover, and what “propagation” really is

By now the interesting work is done and the cutover itself should be dull. It is also the moment to be honest about the phrase in the title of this article. Zero downtime is the wrong target. It is achievable, at the cost of database replication, session sharing and a dual-write period — machinery that is entirely justified when an outage costs more per minute than the machinery costs per year, and entirely unjustified below that line. For most sites the right target is not zero. It is a window that is short, bounded, and chosen by you, scheduled at your quietest hour, announced in advance if anyone would notice. Five minutes at 04:00 on a Tuesday, planned, is a better outcome than a heroic seamless migration that goes sideways at lunchtime, and it is available to everybody.

The sequence, once you have accepted that:

  1. Put the old site into read-only mode — or into maintenance, if it has no read-only mode. This is the start of your window, and it is the only point where the clock is running.
  2. Final differential copy of files, then dump and load the database. Minutes, because the bulk moved days ago.
  3. Verify on the new machine through the hosts-file entry: load the site, log in, make one write, read it back. Not a status page — a real transaction.
  4. Change the record. One A record, one AAAA record, the short TTL you published yesterday. This is the moment the window ends for each visitor individually, as their resolver’s copy expires.
  5. Leave the old server running, still read-only. Do not stop it, do not cancel it, do not repurpose it. It is your rollback and your instrument panel.

What you are watching for now is not a percentage on a propagation checker. Those tools query a list of public resolvers and tell you what those resolvers currently hold, which is interesting and not authoritative about your actual visitors. The real instrument is the old server’s access log. Every request still arriving there is a client whose resolver has not yet re-asked, or who has your address hard-coded somewhere. The rate falls off along the curve your old TTL dictated, and when it goes quiet, the migration is over — for real, and measurably, rather than because a countdown finished.

Expect the tail to be longer than the TTL says. A small number of clients cache aggressively, some networks clamp low TTLs upward, and a few integrations resolve once at boot and never again. Hours, not days, if you lowered the TTL properly. Days, if you did not.

The first hour, and the week after

The first hour is for the things that only reveal themselves under real traffic, and they arrive in a predictable order. Static pages work immediately, which proves very little. Then come the failures of the second inventory — the webhook that cannot reach you, the upstream that rejects your new address, the job that did not run. Watch the new server’s error log rather than its status page, and watch the old one’s access log at the same time.

Mail is the exception that deserves its own paragraph, because it fails on a delay long enough to look like something else. Your new address has no sending history, which means it is not trusted and not distrusted — it is unknown, and unknown is treated as suspicious by design. Add the new address to your SPF record before the cutover rather than after: SPF is additive, both servers may legitimately send during the overlap, and a message from an address the record does not list is a message you have personally instructed the receiver to distrust. Keep the same DKIM selector and key so signatures stay valid across the move. Set the PTR in your new provider’s panel on the day the address is assigned, not on the day mail starts bouncing. What a fresh address is actually worth, and how long it takes to mean anything, is a subject in itself — the point here is only that the cutover is what creates the situation.

Then the week after, which has exactly three items and no urgency:

  • Put the TTL back up once the tail is quiet. A permanent 300-second TTL is not free: it multiplies queries against your zone, and it means a DNS outage becomes a site outage in five minutes instead of an hour. The short value is a tool for a migration, not a configuration.
  • Turn the old server off before you cancel it, and leave it off for a week while still paying for it. A stopped machine you can start again is a backup; a cancelled machine is a decision. The gap between the two costs one month of the smallest plan on offer.
  • Take a backup from the new machine and restore it somewhere. Not verify — restore. You have just changed every assumption your backup process was built on, and the first restore after a migration is the one that discovers the backup was writing to a path that no longer exists.

What is different when nobody is doing it for you

Everything above applies to any two hosts. Four things change when the destination is a provider in another jurisdiction that deliberately holds very little about you, and it is better to know them before the hosts-file entry than after.

There is no transfer team, and this is structural rather than stingy. Mainstream hosts offer to migrate your site for you; that offer requires handing them credentials to the machine you are leaving, and it requires a company willing to hold an account that ties both ends together. A provider whose entire proposition is that it collects and keeps as little as possible is not going to be the one logging into your old server. The work in this article is yours. The compensation is that nothing about the move is contingent on someone else’s queue.

Distance changes how you copy, not how much. Bulk transfer is bounded by bandwidth and is fine — a large tree over a good link is a long lunch, not a project. What round-trip time punishes is many small files, because per-file handshakes multiply by the latency rather than by the size. Moving tens of thousands of small files across a few dozen milliseconds is dramatically slower than moving one archive of the same bytes. Stream an archive for the first pass and use a differential sync for the second; that one decision is usually the difference between an afternoon and a weekend.

Your users’ latency changes, and you should measure it rather than assume it. The number that matters is not your own ping from wherever you happen to be sitting. It is the round trip from the regions your traffic actually comes from, which you can read off your own logs and check against a published latency map before you commit to anything. For most applications a few dozen milliseconds is invisible. For a chatty single-page application making twenty sequential requests to render a view, it is not — and the fix is to stop making twenty sequential requests, which is worth doing regardless of where the server lives.

The relationship you are leaving does not get a grace period. Plan on the assumption that the old provider stops being helpful the moment you give notice, and that anything of yours in their control panel — the zone, the domain, a backup you never downloaded — is easier to retrieve today than next week. This is the same reason the zone moves first, and it is the one part of the sequence that has nothing to do with technology.

One last thing, and it belongs at the end rather than the beginning. Everything here describes how to move well. Whether to move at all is a different question with a different answer, and a migration undertaken to fix a vague unease usually gets unwound within the year at twice the cost. If you have not settled that part, the five cases where a jurisdiction change buys you nothing is the better use of the next ten minutes. If you have settled it, lower your TTL today and read this again tomorrow.

Written by the engineers who run the platform, and re-read 23 days ago. If something here is wrong or has gone out of date, say so from the panel — that is where about half of these came from.

Language

Read this site in your language