Or, "How I finally exited my position in AWS"

GitHub repository: https://github.com/mashiox/s3ql-time-machine/

Background and Motivation

A few years ago, before I fully understood the operational lock-in and caveats that come with S3QL, I deployed several encrypted, deduplicated filesystems directly on Amazon S3. The intent at the time made sense: use S3QL as a secure networked filesystem in the cloud to store sensitive disk backups, and data.

AWS is an exceptional platform. The reliability, API consistency, are second to none. However, AWS is extremely expensive for long-term storage and data egress. As data sizes grew into hundreds of gigabytes and terabytes, keeping static media files in S3 buckets across multiple regions became financially unreasonable.

My primary motivation was straightforward: change the hosting strategy to something more economic. I decided to transition my data off AWS and bring it into a local, self-hosted environment. That was where I ran into a my first blocker.

The S3QL Version Lock-in Problem

On my workstation running Arch Linux (btw), the native S3QL package is on version 6.2.2. When I first attempted to mount one of the legacy buckets, the tool failed immediately:

$ mount.s3ql s3://eu-central-1/eur-media /mnt/s3ql/eur-media/
ERROR: File system revision too old, please run `s3qladm upgrade` first.

S3QL breaks file data into compressed, deduplicated chunks and coordinates all filesystem structures through a central SQLite metadata database stored in S3 as an object named s3ql_metadata. Across major releases, S3QL frequently changes its internal database schema and metadata representation. Newer binaries strictly refuse to mount older filesystem revisions.

When I tried running the suggested upgrade command, I was blocked again:

$ s3qladm upgrade s3://eu-central-1/eur-media
To upgrade the filesystem, first download file system metadata using the previous
version of S3QL (eg. by running fsck.s3ql)

This created a classic chicken-and-egg problem. To upgrade the filesystem to work with modern S3QL, I needed the historical S3QL binaries to run fsck. Moreover, executing an irreversible schema migration directly against production S3 buckets is high-risk. A failed upgrade can permanently corrupt the metadata, turning encrypted data blocks into unrecoverable garbage.

I established clear constraints for this recovery operation:

  • Non-destructive changes are mandatory.
  • Do not downgrade or alter native packages on the host operating system.
  • Do not mutate remote S3 objects into an unrecoverable state.
  • Enable AWS S3 Bucket Versioning across all buckets as a fallback safety net.

The S3QL Time Machine Architecture

Instead of fighting host dependencies or risking an in-place database upgrade on AWS, I designed the "S3QL Time Machine" approach: run the required historical versions of S3QL inside isolated Docker containers.

The container environment:

  • Base image: Ubuntu 20.04 (providing Python 3.8 and stable compilation toolchains).
  • Installed packages: fuse3, libfuse3-dev, cython3, sqlite3, libsqlite3-dev, cryptsetup, and psmisc.
  • Python dependencies: apsw, pyfuse3, dugong, cryptography, requests, and google-auth.
  • Target S3QL version: S3QL 3.8.1 (Filesystem Revision 24), compiled directly from release sources.

To allow the container to provide a FUSE mount back to the host, the container is launched with:

  • --device /dev/fuse
  • --cap-add SYS_ADMIN
  • --security-opt apparmor:unconfined
  • -v $AUTHINFO_FILE:/root/.s3ql/authinfo2:ro
  • -v $(pwd)/pool:/pool:shared

The shared volume propagation ensures that when mount.s3ql attaches the remote filesystem inside the container to /pool/<bucket>, the unencrypted files immediately appear on the host system at ./pool/<bucket>. From there, I run rsync to copy the data directly to local disk.

Technical Obstacles Encountered and Debugged

During the recovery process across the different buckets, several subtle technical issues arose that required diagnosis.

A. S3 Object Metadata and Version Identification

To inspect bucket metadata without mounting, I wrote a Python inspection tool using boto3. Modern S3QL formats metadata using a custom raw2 format split across x-amz-meta-NNN HTTP headers. boto3 automatically strips the x-amz-meta- prefix, exposing keys as 000, 001, format, etc. While cleartext inspection easily confirms format_version: 2, encryption: AES_v2, and `compression: LZMA, for encrypted buckets S3QL intentionally hides the actual revision number inside the encrypted data attribute to prevent metadata leakage. This confirmed that spinning up a legacy container was the only way to inspect the internal revision.

B. Python ConfigParser Passphrase Syntax

When mounting another bucket, mount.s3ql crashed during argument parsing:

configparser.InterpolationSyntaxError: '%' must be followed by '%' or '(', found: '%B70nUcwE'

S3QL uses Python's configparser with BasicInterpolation enabled by default. Because the passphrase contained a literal percent sign (%), configparser tried to parse it as an interpolation variable. The solution under INI specifications is to escape the percent character by doubling it (%%) in ~/.s3ql/authinfo2.

C. Remote Stale Locks (Exit Code 31)

When attempting mounts on buckets that were previously aborted or not cleanly unmounted in 2020, mount.s3ql halted:

ERROR: Backend reports that fs is still mounted elsewhere, aborting.

S3QL tracks an is_mounted boolean flag in the remote metadata to prevent concurrent writers from corrupting the filesystem. Because earlier sessions died without a clean unmount, is_mounted remained True in S3.

Attempting to mount failed. Running modern fsck.s3ql on the host was out of the question. The fix was to execute fsck.s3ql inside the Time Machine container using the --force-remote flag:

docker exec s3ql_mount_<bucket> fsck.s3ql --log none --force-remote --authfile /root/.s3ql/authinfo2 s3://<region>/<bucket>

This scanned the SQLite metadata tables, corrected inode size inconsistencies, cleared the stale lock, and uploaded clean metadata back to S3, allowing subsequent mounts to proceed.

D. FUSE Ownership and Extraction Permissions

Once mounted, running rsync from the host hit Permission denied errors on directories and files. FUSE mounts default to restricting access to the user that mounted the filesystem (root inside the container). Adding the --allow-other flag to mount.s3ql opened access to the host user, and adding --exclude 'lost+found' to the rsync command allowed extraction of the actual data without stumbling on root-only system directories.

Outcome and Conclusion

Using this containerized isolation technique, I successfully recovered and extracted 4 buckets of encrypted data, totaling just over a 1TB footprint in S3.

The objective has been achieved. Investing the time to reverse-engineer the versioning constraints and wrapping the legacy runtime with Docker, I unlocked all of my archived data without paying ongoing AWS storage fees, without risking metadata corruption, and without destabilizing my host operating system.

However, AWS did manage to ding me one last time, charging me for the exfil bandwith!

AI Use Disclosure

An agentic coding harness was used in the creation of the s3ql time machine code. Architectual decision records, and deep research artifacts can be found in the repository's docs/ directory.