Who We Worked With
Our client operates a social media management SaaS platform with publishing, scheduling, analytics, and client-facing workflows running in Docker containers on a single production server. The stack was managed through Dokploy and relied on external systems for data storage, caching, and private connectivity.
The objective was to create an AWS-based environment where the team could test application changes against a replica of its existing infrastructure without putting live services at unnecessary risk.
The Problem: Replicating Production Without Disrupting It
The client’s application was concentrated on a single server. Although its services were containerized, there was no equivalent AWS environment available for controlled testing. Rebuilding the stack manually would have introduced another challenge: reproducing the operating system, container configuration, service dependencies, and network behavior consistently.
We used AWS Application Migration Service (MGN) to replicate the production server into Amazon EC2. This preserved the existing machine state as a starting point, avoiding the need to rebuild the entire environment manually.
However, a successful EC2 launch proved that the disk could be reproduced, not that the resulting machine was safe to connect to the existing environment.
The production server used two networking mechanisms:
- Tailscale provided private connectivity to internal services.
- Cloudflare Tunnel provided public application ingress.
The replica needed to run the application without reusing production's machine identity or introducing unintended connections to live dependencies. That required validating container health, investigating inherited network state, and testing isolation through actual connectivity.
Building and Validating the AWS Replica
After launching the EC2 instance, we inspected the Docker stack rather than treating successful disk replication as proof of application readiness. Initial testing exposed service-image and availability issues that required further investigation and remediation with the client's team.
We then addressed the network boundaries. Cloudflare Tunnel remained disabled on the test instance to prevent competing public ingress. Tailscale required separate investigation because the replica had inherited local state from production.
To restrict access to live dependencies, we examined network connections from running containers and implemented a host-level DOCKER-USER iptables rule blocking access to the production Redis endpoint. We tested connectivity and inspected firewall counters to verify that the control was being enforced.
This moved the work beyond launching a replica: we were checking whether the application could run in the test environment without retaining unintended access to production resources.
The Incident: When the Clone Inherited Production's Identity
The most consequential issue emerged during Tailscale troubleshooting.
AWS MGN had copied tailscaled.state along with the production server's filesystem. This file contained local Tailscale state associated with the source machine. The EC2 replica consequently retained the source machine's Tailscale identity instead of registering as an independent node.
Our initial attempt to resolve the conflict was to change the clone's hostname. That changed its displayed name but did not replace the persisted identity.
The consequences became clear when we ran tailscale logout on the replica. Because the identity was shared, deauthorizing it also disconnected the original production server from Tailscale. Internal services dependent on that private-network connection temporarily lost access to a backend database.
The impact was confined to private connectivity: the customer-facing application remained available through its separate ingress path. Nevertheless, the incident demonstrated how a test-environment operation could affect shared production infrastructure when machine identity had not been separated.
We resolved the identity conflict by removing the cloned Tailscale state file before bringing the service back up, allowing the EC2 instance to establish a fresh identity and register as a separate device. We applied the remediation to the test instance and documented the procedure for future replicas.
The key distinction was between a machine's displayed name and its underlying identity. Renaming the clone could not resolve a conflict that persisted in its local Tailscale state.
What We Verified, Not Just What We Assumed
Verified outcomes
- Tailscale identity: Identified the shared-identity conflict, applied the fresh-identity remediation to the test instance, and documented the procedure.
- Production isolation: Disabled Tailscale and Cloudflare Tunnel while the test instance was isolated and implemented a
DOCKER-USERiptables rule blocking access to production Redis. - Network enforcement: Tested connectivity from running containers and inspected firewall counters. During the documented checks, rebuilt services connected to the intended test Redis and PostgreSQL endpoints, with no production-resource connections observed.
- Public frontend: Verified the main frontend end to end through an AWS Application Load Balancer routing through Traefik.
Related migration work: PostgreSQL to Amazon RDS
In a separate workstream, we migrated the self-hosted PostgreSQL database to Amazon RDS using AWS Database Migration Service (DMS).
Row counts matched across all 66 tables, and change data capture (CDC) was active with approximately zero recorded replication latency at final verification.
These results validate the database-migration work independently.
Where We Left Things
The engagement established an AWS dev environment and a repeatable approach to validating a replica of the client's existing infrastructure.
The work included:
- An EC2-based dev replica created through AWS MGN, with service-level validation and remediation beyond the initial disk replication.
- A documented procedure for generating a fresh Tailscale identity on a cloned instance, addressing the specific failure mode that disconnected the original machine.
- Network-dependency findings based on observed connections rather than assumptions drawn from service names.
- Isolation controls for the dev application instance, including disabled public tunneling, disabled Tailscale connectivity while isolated, and host-level blocking of the production Redis endpoint.
- An AWS Application Load Balancer with an end-to-end verified route to the main frontend, with multi-subdomain routing identified as a separate configuration requirement.
- A subsequent PostgreSQL-to-RDS migration with row-count validation across all 66 tables and CDC confirmed active at the time of verification.
Why This Approach Worked
The difficult part of replicating a production server is not copying its disk. It is understanding which parts of that disk represent portable application state, which represent machine-specific identity, and which connections could create unintended effects outside the dev environment.
AWS MGN provided a practical way to reproduce the existing server, but it could not determine whether every inherited credential, tunnel, service dependency, or network identity was appropriate for a second machine. That required explicit engineering and validation.
At Metasips, infrastructure migration means more than reproducing a machine. It means validating the resulting environment, understanding which dependencies it can reach, and distinguishing confirmed outcomes from work that still needs verification.
Wondering whether your own migration plan accounts for the identities, credentials, background jobs, and network dependencies that a server clone can inherit? Metasips helps engineering teams build and validate AWS environments with the operational checks that make migrations safer.
Facing a similar infrastructure challenge? Let's talk through your migration and testing requirements.
