Skip to main content
ActivePapers

Some links here are partner links — we may earn a commission if you buy, at no extra cost to you. Details.

Best Data Archiving Software: Top Picks Compared

Data archiving software spans four distinct categories: backup tools, hierarchical storage management (HSM), research data repositories, and enterprise archiving platforms, each solving a different problem. The right choice depends on whether you need point-in-time recovery, long-term bit-level preservation, or regulatory retention. The decisive criteria for research groups are format openness, checksum verification, metadata capture and cost at 10 TB+ – not brand recognition.

  • Archiving and backup solve different problems: backup restores a recent state after a failure; archiving maintains a specific, unchanging version for years or decades.
  • For research groups, the decisive criteria are format openness, checksum verification, metadata capture and cost at 10 TB+ – not brand recognition.
  • Open-source software tools (BagIt, DVC, iRODS, Archivematica, restic, Borg) cover most scientific archiving needs without license fees, but require operational effort.
  • Enterprise platforms (IBM, Dell EMC, Veritas, Commvault) add policy engines, legal hold, and compliance reporting that labs rarely need.
  • SAP’s archiving module (SAP ArchiveLink / Data Archiving, now part of SAP Information Lifecycle Management) is a database offload mechanism, not a general-purpose file archive.
  • Test the restore before trusting an archive; an unverified archive is a hypothesis, not a guarantee.

What Is Data Backup Software?

Data backup software creates a recoverable copy of the current or recent state of a system so that a failure (disk death, ransomware, accidental deletion) can be undone. Backup tools optimize recovery time and recovery point: how quickly you can restore and how much data you can afford to lose. Products like Veeam, Bacula, restic and BorgBackup fit this definition. Backups are typically cyclical: the same data is copied repeatedly, with old copies expiring on a retention schedule.

The distinction is important because backup and archiving have opposite economic aspects. Backup storage is expensive per byte because it must support fast random reads and frequent rewrites. Archive storage is cheap per byte because it is written once and read infrequently: tape, object storage with cold tiers, or optical media. A backup that lasts ten years is not an archive; it is a backup with a long retention window and will inherit every silent corruption accumulated by the source system.

For computational scientists, the practical implication is that a nightly rsync to a lab NAS constitutes a backup, not archiving. If a simulation input file was silently corrupted three years ago and every subsequent backup has copied the corrupted version, no backup generation will save you. An archive with per-file checksums and write-once semantics will.

What Is Archiving Data?

Data archiving means moving data that is no longer actively used to a separate long-term storage where it is kept intact and can be retrieved later. The determining properties are immutability (the archived object does not change), integrity verification (checksums or fixity checks) and retention policy (how long and what triggers deletion). Archiving is a lifecycle decision: data moves from hot storage, to warm, to cold storage, and finally to disposal or permanent preservation.

Standards and conventions anchor this work. The Open Archival Information System (OAIS) reference model, published as ISO 14721, defines the functional entities (ingest, archival storage, data management, access) that any serious archive implements.

Related: — Interactive Python and data-science courses you code directly in the browser.

The BagIt specification (RFC 8493) defines a simple, self-describing directory structure for transferring arbitrary files with manifests and checksums. The FAIR (Findable, Accessible, Interoperable, Reusable) principles govern how research archives should expose metadata. The tools that implement these standards interoperate; tools that invent their own container format create lock-in.

Three archiving models dominate scientific practice:

  1. File Level Archive: A directory tree along with manifests and checksums, stored on tape or object storage. BagIt and tar + sha256sum are the minimum versions.
  2. Repository Archive — a managed service (Zenodo, Dryad, institutional repositories, Dataverse) that assigns DOIs, enforces metadata schemas, and handles preservation.
  3. Application Archive: Data exported from a live system in a read-only form, such as SAP’s data archiving or a database’s cold storage tier.

Each model has a different failure mode. File-level archives fail due to bit rot and loss of documentation. Repository archives fail due to policy changes and funding losses. Application archives fail due to vendor format changes.

Worth a look: — One subscription for university-backed Python and data-science certificates.

What Is Data Archiving in SAP?

Data archiving in SAP involves deleting completed business documents and master data from the active database to archive files, while keeping them readable from within SAP transactions. The mechanism is delivered by SAP Data Archiving (transaction SARA and associated archiving objects) and SAP ArchiveLink, which links archived documents to the applications that display them. In current SAP releases, this functionality is under SAP Information Lifecycle Management (ILM), which adds retention rules, legal hold, and destruction management.

The motivation is performance and cost. An SAP ERP or S/4HANA production database that accumulates years of purchase orders, material documents, and change logs grows until queries and backups slow down. Archiving writes these records to archive files – historically on optical media or tape, now often on content repositories or cloud object storage – and deletes them from the database. The records remain accessible through the same transactions, so users see no functional difference.

SAP archiving is therefore a specialized archive coupled with an application, and not a general file archive. This is worth understanding even outside of SAP shops, because it illustrates the general pattern: archive by moving cold records out of the transactional store, maintain a pointer and reader, and enforce retention centrally. If your group uses SAP for procurement or HR alongside your research computing, the two archiving strategies should share the storage policy but not the tooling.

Comparison Table: Archiving Tools by Use Case

ToolCategoryLicenseBest forWatch out for
BagIt (Library of Congress)File-level packagingPublic domain / open sourcePackaging datasets for deposit or tapeNo built-in storage or retrieval UI
ArchivematicaDigital preservation pipelineOpen source (AGPL)Libraries, archives, long-term OAIS complianceHeavy stack (Docker, Elasticsearch); steep setup
iRODSData grid / policy-driven archiveOpen source (BSD)Multi-site research data with metadata rulesRequires sysadmin commitment
DVCData version controlOpen source (Apache 2.0)Versioning datasets alongside Git codeNot a preservation system; remote is your responsibility
restic / BorgBackupBackup with dedupOpen source (BSD)Encrypted, deduplicated snapshotsBackup semantics, not retention-grade archive
Zenodo / Dryad / DataverseRepositoryHosted serviceCitable, DOI-backed dataset publicationSize limits; metadata schema constraints
IBM Storage Archive / ILMEnterprise archiveCommercialRegulated retention, legal hold, tieringCost and complexity for small labs
Veritas Enterprise VaultEnterprise archiveCommercialEmail and file compliance archivingLicensing model; migration friction
CommvaultBackup + archive platformCommercialUnified policy across large estatesOverkill below a few hundred TB

When selecting data archiving software, users often choose between commercial products and open-source software. Depending on the environment, one might look for open source pc software or specific software in open source communities. While IBM provides enterprise tools, they also contribute to ibm open source software initiatives. For those exploring the history of the field, they may look into the first open source software examples that paved the way for today’s tools.

How to Choose: Criteria That Actually Matter

Opening the format. An archive that you will not be able to read in twenty years is a handicap. Prefer simple files, documented schematics and standard containers (BagIt, tar, Parquet, HDF5 with published specifications) to proprietary bundles. HDF5 and NetCDF are self-describing and widely supported, which is why they dominate simulation results.

Integrity checking. Require checksums at write time and periodic fixity checks. sha256 manifests, BagIt’s manifest-sha256.txt, and repository-side fixity services are all eligible. Without verification, silent corruption is undetectable until you need the data.

Related: — A deep technical library of scientific-computing books, videos and live training.

Metadata capture. An unprovenanced archive is a pile of bytes. Capture the software version, input parameters, random seeds, and environment (container image digest or conda lockfile) alongside the data. This is where research archives differ from IT archives.

Cost at scale. Tape remains the cheapest per terabyte medium for cold data; Cold tiers of the cloud (e.g., object storage archive classes) trade a higher cost per byte for zero operational overhead. Model your cost at 10 TB and 100 TB, not 1 TB.

Retention and deletion. Explicitly decide what will be deleted and when. GDPR-style erasure requests and funder retention mandates (often 5–10 years, sometimes longer) both apply to research data; a policy engine helps.

Reader favorite: — Project-based data-science paths with a guided terminal and real datasets.

Test Restoration. Schedule a restoration exercise. Restoring a random sample of files and verifying checksums is the only proof that the archive is working.

Open-Source Options and Their Place

Open-source software dominates scientific archiving for good reason: the source is verifiable, the formats are documented, and there is no per-seat cost. The open source model itself has a long history – the term was coined in 1998, and the practice of source code sharing predates this by decades, with early examples like the GNU Project (1983) and, before that, academic and ARPANET-era code sharing. Today, a research group’s common open source software stack typically includes Linux, Python, Git, and a package manager, with archiving tools layered on top of it.

For a lab, a workable open source archive might be: DVC or Git LFS for versioning active datasets, BagIt for packaging frozen datasets, restic or BorgBackup for off-site encrypted copies, and Zenodo for DOI-backed publication of the subset that needs to be citable. iRODS or Archivematica come into play when you need policy-driven tiering or formal OAIS compliance across multiple sites. Note that “open source PC software” for archiving is thinner than for other categories: the most mature archiving tools are server-side or CLI-first, and desktop GUI options are often commercial interfaces for open source engines.

One caveat needs to be clearly stated: open source shifts the cost from licenses to labor. A group without a research software engineer should put more weight on hosted repositories, because an unmaintained iRODS instance is worse than a paid service.

Sources & Further Reading

  • Research data archiving — Wikipedia: Research data archiving is the long-term storage of scholarly research data, including the natural sciences, social sciences, and life sciences. The various academic…
  • Open-source software — Wikipedia: Open-source software (OSS) is computer software whose source code is publicly available, allowing users to use, study, modify, and distribute it — in contrast with…
  • Open source — Wikipedia: Open source is the practice of publishing digital resources publicly alongside their source code or source files, enabling use, study, modification, and redistribution…
  • Study software — Wikipedia: Study software refers to computer programs designed to enhance the effectiveness of learning by improving the way students engage with, process, and retain information…

Frequently Asked Questions

What is data backup software?

Data backup software copies the current state of a system so that it can be restored after a failure, corruption, or deletion. It optimizes for fast recovery and short retention windows, and typically rewrites the same data repeatedly. Examples include Veeam, Bacula, restic and BorgBackup. Backup is a recovery mechanism and not a preservation mechanism.

What is archiving data?

Data archiving involves moving inactive data to a separate long-term store where it is kept immutable, integrity-verified, and recoverable under a retention policy. Archives follow standards such as OAIS (ISO 14721) and BagIt (RFC 8493) and prioritize durability over speed of access. The goal is to preserve a specific version and not to recover the latest state.

What is data archiving in SAP?

Data archiving in SAP removes completed business documents from the live database into archive files while keeping them readable through SAP transactions. It is implemented via SAP Data Archiving (SARA transaction), ArchiveLink and, in current versions, SAP Information Lifecycle Management for retention and legal hold. The goal is database performance and cost control, not general file preservation.

Is open-source archiving software good enough for research data?

Open source tools such as BagIt, Archivematica, iRODS, DVC, restic and BorgBackup are used in production by libraries, archives and research computing groups. They meet preservation standards and avoid licensing costs, but require operational expertise. Groups without dedicated systems staff often combine open source packaging tools with a hosted repository for final deposit.

How is archiving different from backup?

Backup restores a recent state after a failure; archiving keeps a specific version for years or decades. Backup storage must be fast and is rewritten frequently; Archive storage is cheap, write-once, and rarely read. A long-running backup is not an archive because it propagates silent corruption forward rather than freezing a verified copy.

How often should I verify an archive?

Check frequency depends on the media and risk tolerance, but a common practice is to run fixity checks on a scheduled basis (quarterly or annually for tape and object storage) and always after media migration. Each audit should compare stored checksums with recalculated ones and record mismatches for investigation. Restoration drills should accompany verification and not replace it.

P.S. A few readers have asked which guided learning paths we actually reach for — it's Dataquest; if you want the current details.

Frequently asked questions

What is data backup software?

Data backup software copies the current state of a system so that it can be restored after a failure, corruption, or deletion. It optimizes for fast recovery and short retention windows, and typically rewrites the same data repeatedly. Examples include Veeam, Bacula, restic and BorgBackup. Backup is a recovery mechanism and not a preservation mechanism.

What is archiving data?

Data archiving involves moving inactive data to a separate long-term store where it is kept immutable, integrity-verified, and recoverable under a retention policy. Archives follow standards such as OAIS (ISO 14721) and BagIt (RFC 8493) and prioritize durability over speed of access. The goal is to preserve a specific version and not to recover the latest state.

What is data archiving in SAP?

Data archiving in SAP removes completed business documents from the live database into archive files while keeping them readable through SAP transactions. It is implemented via SAP Data Archiving (SARA transaction), ArchiveLink and, in current versions, SAP Information Lifecycle Management for retention and legal hold. The goal is database performance and cost control, not general file preservation.

Is open-source archiving software good enough for research data?

Open source tools such as BagIt, Archivematica, iRODS, DVC, restic and BorgBackup are used in production by libraries, archives and research computing groups. They meet preservation standards and avoid licensing costs, but require operational expertise. Groups without dedicated systems staff often combine open source packaging tools with a hosted repository for final deposit.

How is archiving different from backup?

Backup restores a recent state after a failure; archiving keeps a specific version for years or decades. Backup storage must be fast and is rewritten frequently; Archive storage is cheap, write-once, and rarely read. A long-running backup is not an archive because it propagates silent corruption forward rather than freezing a verified copy.

How often should I verify an archive?

Check frequency depends on the media and risk tolerance, but a common practice is to run fixity checks on a scheduled basis (quarterly or annually for tape and object storage) and always after media migration. Each audit should compare stored checksums with recalculated ones and record mismatches for investigation. Restoration drills should accompany verification and not replace it.


Build a data portfolio, project by project

Project-based data-science paths with a guided terminal and real datasets