Skip to main content

The Data Is Still There. I Have Already Lost It.

Robert
Data SovereigntyPersonal DataSelf-custodyLocal-firstArcSpaceAI

Editor's note

ArcSpace is approaching a new release. Its predecessor is DID Space, familiar to many ArcBlock users: a personal data space tied to a user's own identity, with application access controlled by that user. Before the release, ArcBlock founder Robert steps back to ask what “users own their data” must actually mean, drawing on more than twenty years of disks, video tapes, cloud records, and private keys.

I have an MO disk from 2001, a rewritable magneto-optical disk of the kind people used back then. The label says “Gold Backup.”

I believe the data is still on it. The disk was not lost or stolen. No company shut it down. I kept it carefully. There is only one problem: I no longer own a drive that can read it.

So do I still own that data?

We talk a lot about data ownership, self-custody, and personal data sovereignty, as if moving files from somebody else's cloud back onto our own hard drive settles the question. I have come to believe something harder: real data sovereignty is not about where the data sits today. It means that after devices, formats, applications, providers, and even your own memory have changed, you can still access, understand, migrate, verify, and recover the data, and decide when to delete it.

This is not a choice between decentralization and centralization. My own experience happens to include the best and worst of both.

Put the story on a timeline first

The rest of this story can be compressed into one reading map. Each generation of technology solved a real problem from the generation before it. It also moved control and the point of failure somewhere new.

EraWhat we gainedThe new vulnerability
Personal media: floppy disks, tapes, optical discs, hard drivesFiles in our own hands, available offlineMedia fail; formats and readers disappear
Internet services: email, forums, photo sites, notesImmediate access, search, sharing, and use across devicesServices close, change owners, start charging, or make accounts unrecoverable
Cloud platforms: managed databases and object storageProfessional backup, migration, indexing, and availabilityThe provider's record becomes the final truth; independent proof and recovery are difficult
Crypto: private keys, smart contracts, and on-chain assetsFinal signing authority and publicly verifiable recordsKeys are lost; hardware and contracts fail; exchanges, DeFi, and bridges are attacked
Local-first software and personal data spacesLocal control and online collaboration no longer need to be oppositesAuthorization, sync, recovery, and format migration still require long-term maintenance
AI and ArcSpaceOrdinary people may be able to organize, migrate, and inspect their data continuouslyAI must remain constrained by identity, permission, and auditable rules instead of becoming a new single point of power

This is not a straight line from bad to good. It is a history of answering the same question again and again. Location is not enough. Who can keep using, moving, proving, recovering, and deleting the data determines who truly owns it.

A room full of “ownership”

I started making backups early. Five-and-a-quarter-inch floppies, three-and-a-half-inch floppies, CDs, DVDs, MO disks, portable hard drives, SSDs, and NAS systems: I have kept some version of nearly all of them. I also have dozens of Mini DV tapes containing family life, company meetings, and whatever else seemed worth recording at the time.

A small media glossary

  • MO, or magneto-optical disk, writes data with a combination of lasers and magnetic fields. It was valued for rewritability and durability, but working readers are now uncommon.
  • Mini DV stores digital video on small magnetic tapes. The recording may survive, but migration requires a camera, FireWire, and the rest of the period's playback chain.
  • SSD stores data in flash memory. With no spinning platter it is fast and resistant to shock, but charge retention, controllers, interfaces, firmware, and write wear still matter.
  • NAS is storage connected to a home or office network. It can share files across devices and provide redundancy, but it still requires maintenance and is not automatically an off-site backup.

A box of Mini DV tapes kept for decades. One label still says 1997.

Most of these things are still physically mine. But “still there” is more complicated than it sounds.

Mini DV is slightly better because I still have a portable camera that can play the tapes, barely. Finding equipment with FireWire that can transfer dozens of tapes properly is another matter. The MO disk may not have lost a single bit, but I cannot even test that claim. I once assumed that a hard drive put away safely was safe forever. Then I encountered drives that remained on the shelf while their data became unreadable. SSDs taught another lesson: having no moving parts does not make a device suitable for indefinite unpowered storage. Flash memory holds data as electrical charge, and retention depends on temperature, wear, and the type of flash. An SSD can be very reliable. It is not a stone tablet you put in a drawer and forget for twenty years.[1]

This SanDisk portable SSD looks intact but no longer works. Possessing the device does not make its data reachable.

An SSD is not a permanent cold archive

One baseline industry requirement for a consumer SSD is that after reaching its rated write endurance, it should retain data without power for at least a year under specified temperatures. A lightly used drive kept in normal conditions may last much longer. The one-year requirement is still a useful reminder: the design goal was never “leave it in a drawer for decades.” The SSD in the photo stopped working after only two or three years in storage. Its appearance cannot tell us whether the failure was in the flash, circuitry, controller, interface, or firmware. That uncertainty is the point.[1]

During another move, a whole box containing backup drives disappeared with the moving company. Compensation could be negotiated. No compensation could restore the data.

These are hard lessons I paid for myself: owning the medium is not the same as owning usable data.

Archivists understood this long ago. When the U.S. National Archives discussed optical-disc preservation in the 1990s, it broke long-term access into ordinary questions: can a machine read it, can you find it, and can you understand it once found? The report warned that a disc may outlive the hardware and software needed to read it. What must be preserved is access to the information, not devotion to one piece of media.[2]

That observation has not aged at all. A video with no filename, date, people, or context is still only a few gigabytes of pixels, even if it plays. A structurally complete export from an application may be equally useless if ordinary people cannot reconstruct their lives without the original interface and explanations.

Professional digital preservation was never “make one backup.” It is continuous work: check for corruption, keep multiple copies, make the copies independent, and migrate before media and formats become obsolete. Stanford Libraries' LOCKSS program put one principle in its name: Lots of Copies Keep Stuff Safe. More interestingly, it says the copies should not all be controlled by the same system and the same administrator. Otherwise many apparent copies can still disappear together.[3]

Ordinary people cannot run a national archive at home. That is exactly why translating “users own their data” into “put everything on the user's own hard drive” is irresponsible.

The cloud solved real problems and gained real power

The other half of my experience happened on the Internet.

Early BBSes, forums, and technical communities, along with the people I met and the things I wrote there, mostly survive only in memory. I started using Hotmail in the 1990s. Seeing my email from 1996 would be fascinating, but it is long gone. My Flickr photos remain publicly visible, yet after several changes of ownership I can no longer recover the account. The free tier shows only an initial set of photos. I can see that more exist, but I cannot retrieve them. My old Evernote notes are still there too, behind what now feels like a very expensive subscription.

These stories invite an easy conclusion: do not trust the cloud.

That is not my conclusion. Some of my oldest and most searchable personal data lives in Gmail, Google Drive, and Google Docs. It has crossed decades of computers, hard drives, offices, and moves, and remains at my fingertips. Google has handled backup, migration, indexing, and format maintenance better than I could have done alone. I suspect this is true for millions of ordinary users. The cloud is more than “somebody else's computer.” It is professional maintenance that most people cannot reproduce by themselves.

The problem begins when everything is perfect until it is not.

I was an early invited user of Google App Engine and later a heavy paying customer. I kept structured data for blogs, notes, and hobby projects in Datastore. I even praised the service in media interviews because I honestly believed Google's backups must be better than mine.

Then, after a platform upgrade, the data vanished. Official support investigated and told me that their records showed those databases had never existed.

That sentence taught me the boundary of cloud data ownership. I remembered the databases. I had used them and paid for them. But in the provider's system of record, they did not exist. My memory was not evidence, and I had no independent record proving otherwise.

By then I was working in Seattle and knew many people at Google. Some were sympathetic but lacked permission to help. Eventually I reached a former Microsoft colleague working in the relevant organization. He said the data should in principle still exist, but without an official ticket he could not cross internal access rules to search historical records. A real investigation might require a lawyer's letter that forced the company to open a legal process.

Suing Google over a few hobby projects made no sense. I gave up. Since then I have never placed the only copy of important data in Google App Engine or any other cloud. This does not mean Google is uniquely bad. The same thing could happen at any large company, perhaps even in a service we operate. The fact that no villain is required is exactly why the problem matters.

The year 2021 brought a more public reminder. Twitter and Facebook restricted then-U.S. President Donald Trump's accounts under their own rules. AWS later suspended Parler over content moderation and terms-of-service disputes, taking the service offline for a time.[4] People can hold completely different political views about those decisions. Only one point matters here: a platform does not merely store data. It exercises practical control over whether that data can continue to be accessed, distributed, and operated.

The cloud is not the enemy. A cloud service that offers no practical exit, proof, migration, or independent recovery is the problem.

Not all data should live forever

At this point it is easy to run toward the opposite extreme. If loss is so dangerous, preserve everything for as long as possible.

I do not agree with that either.

Data sovereignty includes the right to keep data and the right to forget it. A youthful draft, a private conversation, location history, medical information, and identity material whose purpose has expired should not be retained forever by default simply because storage keeps getting cheaper. The European Union's GDPR contains data portability, a right to erasure, and storage limitation. Under applicable conditions people can request a transfer or deletion, while organizations should not keep personal data longer than its original purpose requires. Historical, scientific, and public archives may have exceptions, with safeguards.[5]

This may seem to conflict with an archive's mission of long-term preservation. It is actually the same question: who has the authority to decide why data exists and how long it should exist?

Some data is work and memory worth carrying across a lifetime. Some is an operational record needed only until a dispute can be resolved. Some is sensitive enough that deletion is safer once its purpose ends. Saving everything forever is not sovereignty. It is surveillance moved onto a different hard drive.

Real control includes retention rules and deletion. It must also accept an uncomfortable fact: deletion is an engineering problem too. If ten invisible backups exist, “deleted” may only mean gone from the interface. If users personally control every copy, they may discover during recovery that they deleted too thoroughly.

The question was never local or cloud. The question is whether users know which copies exist, who can read them, why they are retained, when they expire, whether they can be moved, and who is responsible for recovery.

Crypto turned the same question into an extreme experiment

The loss of ordinary data may go unnoticed for years. In crypto, the same mistake can become an immediate loss of assets.

I have lost private keys more than once. Most were not hacked. I misplaced a backup, forgot a password, or discovered years later that a recovery method I trusted no longer worked. The loss did not always hurt at the time. Years later, when I understood what the key represented, nobody could help.

“Not your keys, not your coins” identifies a real risk of centralized custody. It is often misread as its converse: “Your keys, therefore your coins are safe.”

They are not.

A Ledger Nano X hardware-wallet box. Hardware can protect a key, but it cannot maintain backups, recovery, and long-term access for its owner.

Holding the key means nobody else can unilaterally move the asset for you. It does not guarantee that you will remember the password, retain the seed phrase, or use hardware and software that remain correct forever. The COLDCARD incident disclosed in 2026 is a harsh example. The hardware wallet appeared normal and contained a secure random-number generator, yet an integration error caused some devices to generate private keys through a much weaker random path. Users had chosen self-custody and still suffered serious losses through an implementation they could not personally inspect.[6]

The 2017 Parity multisig incident was different. Users had not placed funds in an exchange or simply written a key on paper. They used a smart contract that required multiple approvals. When shared program code on which the wallets depended was destroyed, funds in 587 wallets became permanently frozen. The keys were neither lost nor stolen, but the assets could no longer move. To the owner, the result was indistinguishable from loss.[7]

NFTs exposed another layer. A token being on-chain does not put the image and metadata it references on-chain. A study of 12,353 NFTs cited in NIST's 2024 NFT security report found that 25 percent already pointed to missing or inaccessible assets.[8] You may permanently own an address that still works while everything behind it has disappeared.

Giving up self-custody does not remove risk

Those examples can create another illusion: avoid self-custody, hand assets to an exchange, DeFi protocol, or blockchain service, and the problem goes away.

It only moves.

Mt. Gox was once one of the world's largest Bitcoin exchanges, with servers holding customer wallets and private keys. U.S. Department of Justice charges made public in 2023 allege that attackers gained access to its servers beginning in 2011 and ultimately stole approximately 647,000 BTC.[9] Customers did not lose their own seed phrases. They made recovery dependent on the exchange's systems, security, and liquidation process.

DeFi and DEX systems remove the traditional account administrator but add boundaries around contracts, oracles, governance, front ends, and cross-chain bridges. A 2022 FBI warning reported roughly $1.3 billion in cryptocurrency stolen during the first quarter of that year, with almost 97 percent taken from DeFi platforms. It listed flash loans, signature-verification flaws, and price-oracle manipulation among the attack paths.[10] The Ronin Bridge was attacked the same year for about $620 million. The U.S. Treasury later identified it as an attack against the blockchain project associated with Axie Infinity; Ronin's own postmortem recorded 173,600 ETH and 25.5 million USDC removed from the bridge.[11]

A CEX, a DEX, a bridge, and “the chain was hacked” are not the same thing

A CEX generally has an operator maintaining accounts and custody. A DEX or DeFi protocol organizes transactions and financial logic primarily through on-chain programs. A bridge represents and transfers assets between chains. Headlines often say a chain was hacked when the failure may instead be in consensus, an application contract, validator keys, an oracle, a front end, or a bridge. Finding the actual layer matters because it tells us who had control, who could repair the damage, and who bore the loss.

Crypto did not prove self-custody wrong, and it did not prove custody safer. It proved that custody is only one part of ownership. A private key grants final authority and assigns you final responsibility for recovery. An exchange takes on some operations while making you depend on its ledger and security. A smart contract moves some trust from a company into code and introduces the risk of program failure. A blockchain can prove that a record exists. It cannot keep a bridge, an application, or off-chain content available.

This is the same story as the MO disk in my hand. The medium, key, token, and file may all remain. The real object can still be unreachable.

Ownership should be a continuing capability

In 2019, Martin Kleppmann and his colleagues at Ink & Switch published Local-first software: You own your data, in spite of the cloud. They did not reject the cloud's collaboration and multi-device access. They asked whether software could provide those benefits together with the ownership of old-fashioned software. Their local-first principles influenced many later products, while candidly acknowledging that ownership also creates responsibilities for backup, ransomware protection, and long-term maintenance.[12]

Solid and Personal Data Stores approach a similar question from another direction. They separate data from applications, allowing users to choose a personal data space and decide which applications may read or write it. Your photos can remain in a space you choose when you switch gallery applications; you do not need to hand the whole collection to another company, and you can revoke the old application's access. Solid also emphasizes standards and interoperable data formats. More than a decade of Personal Data Store research has examined authorization, sharing, migration, privacy, and whether ordinary users can actually operate these systems.[13]

This work matters because it turns “ownership” from a slogan into capabilities we can test.

For me, data that truly belongs to a user should survive a series of questions. Can I access it directly today? Will its format still be understandable in ten years? Can I move it completely to another provider? Can I verify that its contents and history were not silently changed? Can I recover after hardware failure, a forgotten password, or a provider shutdown? Can I see, revoke, and expire permissions? Can I actually delete it when I no longer need it?

Not every piece of data needs the highest standard forever. Some is worth seven days, some seventy years. The point is not to turn every box green. These decisions cannot live only inside a provider policy, an application database, or a drawer the user has forgotten.

A personal data system also cannot demand that everyone become a system administrator. People should not need to study drive-retention limits, file formats, integrity checks, off-site copies, and key rotation before they deserve to own their photos and notes. Data sovereignty that only experts can operate still belongs only to experts.

AI may finally separate “managing it yourself” from “doing everything yourself”

The hardest part of these systems was never the principle. It was the operation.

Today I keep important data locally, maintain another copy elsewhere, and use more than one online provider. I distribute digital assets and place recovery material in different locations. None of this guarantees 100 percent safety. More copies can create more opportunities for attack. At least one drive, one account, one administrator, or one provider is no longer the only point of failure.

A Seagate portable drive still in my possession. Keeping the medium is only one link in a recovery chain.

The problem is that this is tiring, and I will never do it perfectly.

What an AI assistant may change is not storage itself but the cost of maintaining it. It can keep track of which data has only one copy, which drives have not been checked, which formats are losing tool support, and which permissions have expired. It can recommend migrations, build indexes, test whether a backup is genuinely recoverable, and carry out a move when you change providers. More personally, it could turn an unlabeled Mini DV tape into searchable memories with dates, people, and events.

That does not mean “let AI own the data.” Quite the opposite. An AI assistant should work within permissions the user explicitly grants, leave an inspectable record, require confirmation for important deletion and authorization changes, and support reversal. Otherwise we have only replaced the cloud provider's concentrated power with a model that is harder to understand.

These conclusions have also become constraints for our design of ArcSpace. ArcSpace grew from DID Space, and the name change reflects a change in role. The data space remains associated with the user's own DID and applications receive access only as authorized. Within ARC, however, a user-held data space has moved from an add-on service to a foundational capability. Applications should request the permissions and data model required for a task instead of inheriting an entire account.[14] The design must continue to answer more concrete questions: an online provider cannot be the only authority; local and hosted copies must coexist; and every important AI operation must be explicitly authorized and recorded.

ArcSpace is still evolving. It does not yet solve long-term format migration, and it cannot promise that no user will ever lose data. What we want to test is a system that does not force one permanent choice between convenience and control. Users can rely on professional online services without making the service provider the sole authority. They can hold local copies without carrying all maintenance alone. Future AI assistants can help organize and migrate, but cannot override the user's identity and permissions.

I still want to read that 2001 “Gold Backup.” I want to organize those dozens of Mini DV tapes and let AI help me see what is on them. They remind me that data sovereignty is not about gripping an object more tightly. Data that truly belongs to you should stay alive as it travels with you. Devices fail, formats age, applications stop, providers change, and you forget passwords. After all of that, if you can still find the data, understand it, carry it away, or personally decide to let it go, then “yours” means more than a word written on a disk case.

References


  1. Kingston, “Myths and Misinformation About SSDs”. NAND retention varies with temperature, wear, and flash type; an SSD is not an indefinite offline archive. ↩
  2. U.S. National Archives and Records Administration, “Digital-Imaging and Optical Digital Data Disk Storage Systems”. ↩
  3. LOCKSS Program, “Frequently Asked Questions”; “Preservation Principles”. ↩
  4. Twitter, “Permanent suspension of @realDonaldTrump,” January 8, 2021; Oversight Board, decision on Donald Trump's suspension, May 5, 2021; U.S. District Court, Western District of Washington, Parler LLC v. Amazon Web Services Inc., January 21, 2021. These sources also record the platforms' stated reasons and the surrounding disputes. This article takes no position on the moderation decisions themselves. ↩
  5. European Commission, “Principles of personal data processing under the GDPR”; “Information for individuals”. ↩
  6. COLDCARD, “Current COLDCARD Security Status”; HodleHub, “Coldcard: audit review,” August 1, 2026. ↩
  7. Parity Technologies, “A Postmortem on the Parity Multi-Sig Library Self-Destruct,” November 15, 2017. ↩
  8. NIST IR 8472, Non-Fungible Token Security, 2024. ↩
  9. U.S. Department of Justice, “Russian Nationals Charged With Hacking One Cryptocurrency Exchange And Illicitly Operating Another,” June 9, 2023. These are public allegations; defendants are presumed innocent unless and until proven guilty in court. ↩
  10. FBI Internet Crime Complaint Center, “Cyber Criminals Increasingly Exploit Vulnerabilities in Decentralized Finance Platforms,” August 29, 2022. The overall share cited there came from Chainalysis statistics available at the time. ↩
  11. U.S. Department of the Treasury, “U.S. Treasury Issues First-Ever Sanctions on a Virtual Currency Mixer, Targets DPRK Cyber Threats,” May 6, 2022; Ronin Network, “The Ronin Bridge Is Open,” June 28, 2022. ↩
  12. Martin Kleppmann, Adam Wiggins, Peter van Hardenberg, and Mark McGranaghan, “Local-first software: You own your data, in spite of the cloud,” 2019. ↩
  13. Solid Project, “About Solid”; Shahriar Akter et al., “Personal Data Stores (PDS): A Review,” 2023. ↩
  14. ArcSpace; DID and personal data spaces. ↩