Molecular Dynamics Repository: A Practical Guide
A molecular dynamics repository is a version-controlled host for simulation code, input datasets, force field parameters, and analysis scripts, spanning engines like GROMACS, LAMMPS, and OpenMM plus trajectory formats such as DCD and XTC. Storage choices matter at scale: a 100,000-atom system simulated for 100 nanoseconds with coordinates written every 10 picoseconds yields roughly 10,000 frames, so the right structure determines whether a simulation stays reproducible years later.
Key Takeaways
- A molecular dynamics repository is not one thing: it is a stack of engine code, force-field definitions, topology and coordinate files, run scripts, and analysis notebooks, each with different versioning needs.
- Binary trajectory formats (XTC, DCD, TRR, NetCDF/AMBER) trade precision against size; the choice affects both storage cost and long-term readability.
- Reproducibility depends less on the engine than on pinning four things: engine version, force field version, random seed, and the exact input file.
- Git is the right tool for text inputs and scripts, but large binary trajectories belong in Git LFS, DVC, or a data repository such as Zenodo — not in the main Git history.
- FAIR data principles (Findable, Accessible, Interoperable, Reusable) map cleanly onto repository layout decisions you make on day one.
- A repository that cannot be re-run by a stranger is documentation, not reproducibility.
What “Molecular Dynamics Repository” Actually Refers To
Molecular dynamics repository is an overloaded phrase, and the ambiguity causes real confusion when lab groups try to standardize. At least four distinct things get called by that name, and they have almost nothing in common technically.
The first is the engine repository — the source code of the simulation program itself. GROMACS, LAMMPS, NAMD, OpenMM, and AMBER each live in their own upstream repositories, and you interact with them as a user, not a maintainer. Here you care about release tags and version numbers, not the code.
The second is the repository of force fields and parameters. Force fields such as CHARMM36, AMBER ff14SB, OPLS-AA and the coarse-grained Martini family are distributed as parameter files, often with their own versioning. The OpenKIM project maintains an interoperable archive of interatomic potentials for materials and molecular systems, which is a truly different delivery model than a Git repository.
The third is the project repository: your own working directory for a specific study. This contains the topology files, coordinate files, .mdp or input decks, run scripts and analysis code. This is where most reproducibility failures come from.
The fourth is the data repository: a long-term archive for trajectories, checkpoints and derived data. Zenodo, Figshare, and institutional repositories fulfill this role and assign DOIs so that a specific dataset can be cited.
Related: — Interactive Python and data-science courses you code directly in the browser.
Confusing these four leads to predictable mistakes: committing a 40 GB trajectory into Git, or treating a Zenodo deposit as if it were an active working directory. Separating them is the single highest-leverage decision in setting up a simulation project.
The Long Data Problem: Why Trajectories Break Normal Version Control
The long data produced by molecular dynamics simulations constitute the determining constraint on repository design, and it is important to understand why before choosing the tools. A single system of 100,000 atoms simulated for 100 nanoseconds with coordinates written every 10 picoseconds produces on the order of 10,000 frames. Stored as uncompressed double precision coordinates, that is approximately 100,000 atoms × 3 coordinates × 8 bytes × 10,000 frames — hundreds of gigabytes before any compression.
Trajectory formats exist precisely to manage this. XTC uses lossy compression with configurable precision, typically 3 decimal places in nanometers, and is the default for GROMACS. DCD is the classic CHARMM format, uncompressed and widely readable. TRR is GROMACS’ lossless full-precision format, useful when you need exact velocities or forces. NetCDF-based formats, including the AMBER trajectory convention, are self-describing and contain unit metadata.
Reader favorite: — Project-based data-science paths with a guided terminal and real datasets.
The practical consequences for a repository are concrete:
- Git stores every version of every file. A trajectory that changes on every run will bloat a repository permanently, because Git history is append-only. Even deleting the file does not reclaim the space without history rewriting.
- Git LFS (Large File Storage) replaces large files with pointers and stores the content elsewhere, which works well for files up to a few gigabytes that change rarely. It is a poor fit for trajectories that are regenerated constantly.
- DVC (Data Version Control) tracks data by hash and stores it in a configurable remote — local disk, S3, or an institutional store — while keeping small
.dvcpointer files in Git. This suits simulation workflows where the data is large and the code is small. - Data repositories with DOIs are the right home for the frozen, published version of a trajectory. They are not a working directory.
A workable division of labor: Git for scripts, input decks, and analysis code; DVC or Git LFS for moderate binary artifacts; a DOI-issuing archive for the final published dataset. This keeps clone times sane and makes the published record citable.
Anatomy of a Reproducible Simulation Repository
A reproducible molecular dynamics repository has a predictable layout, and the layout itself communicates intent to anyone who opens it. The following structure is a reasonable default for a molecular dynamics project, adaptable to LAMMPS, GROMACS, or OpenMM workflows.
project/
├── README.md
├── environment.yml # or requirements.txt / conda spec
├── systems/
│ ├── system-a/
│ │ ├── topology/
│ │ ├── coordinates/
│ │ └── parameters/
│ └── system-b/
├── simulations/
│ ├── equilibration/
│ │ ├── inputs/
│ │ └── run.sh
│ └── production/
│ ├── inputs/
│ └── run.sh
├── analysis/
│ ├── notebooks/
│ └── scripts/
├── data/ # DVC-tracked or gitignored
└── docs/
Each directory deserves its place. The systems/ tree separates the chemical definition of what you are simulating from how you are simulating it, which is important because the same system is often run under multiple protocols. The simulations/ tree separates equilibration from production, a distinction that is easy to lose and costly to rebuild.
The environment.yml file is the most underrated component. Pinning the engine version, the Python analysis stack, and all supporting libraries into a single file means a collaborator can rebuild the environment with a single command. Conda environment files, pip requirements files, and container definitions (Docker or Apptainer/Singularity) all serve this purpose; containers are the most robust because they also capture system libraries.
The README.md should indicate, at a minimum: what the study is, what engine and version, what force field and version, how to run the pipeline from raw inputs to final figures, and where the large data is located. A README that assumes the reader is the author is not documentation.
Choosing a Repository Host: Criteria That Matter
Repository hosting for computational science falls along lines that commercial Git hosting does not cover. The table below compares the options actually considered by most laboratory groups when selecting a molecular dynamics repository.
| Host / tool | Best for | Handles large binaries | DOI / citation | Notes |
|---|---|---|---|---|
| GitHub / GitLab | Code, scripts, small inputs | Via Git LFS (quota-limited) | No native DOI | Ubiquitous; LFS quotas can surprise you |
| Zenodo | Frozen published datasets | Yes, generous limits | Yes, per-version DOI | Integrates with GitHub releases |
| Figshare | Datasets, figures, supplementary | Yes | Yes | Common in journal workflows |
| Institutional repository | Long-term institutional archiving | Varies | Usually yes | Persistence tied to the institution |
| DVC + cloud remote | Active large-data versioning | Yes | No | Keeps Git history small |
| Open Science Framework | Project-level organization | Yes | Yes | Good for mixed code/data projects |
The decision generally comes down to three questions. Should the data be citable with a DOI? Should it be versioned as it changes, or frozen once? Who is responsible for keeping it available in ten years?
A common and defensible pattern is GitHub for the code with Zenodo integration enabled, so that each tagged release mints a DOI automatically, plus DVC for the working data. This gives citable releases without bloating the Git history.
Force Fields, Parameters, and the Versioning Trap
Force-field versioning is where reproducibility quietly fails, and it deserves separate treatment because the failure mode is invisible. Two simulations titled “CHARMM36” run three years apart may use different parameter sets because the force fields are revised. The same is true for the AMBER protein force fields, where ff99SB, ff99SB-ILDN, ff14SB and later revisions produce measurably different behavior.
The catch is that force field files are often distributed across engine installations or downloaded from a project website, and the version is not recorded anywhere in the simulation output. A molecular dynamics repository that pins the engine version but not the force field version is only half reproducible.
Practical mitigations:
- Vendor the parameter files into the repository. Copy the exact
.itp,.prmor.frcmodfiles used intosystems/*/parameters/and commit them. These are small text files, and their presence in the repository removes any ambiguity. - Save the force field name and revision in the README and in a machine-readable metadata file. A short YAML or JSON file alongside the inputs costs nothing and definitively answers the question.
- Note all local modifications. If you adjusted a partial charge or a bonded parameter, that change should be in the repository, not in a personal scratch directory.
- For interatomic potentials in materials simulation, prefer a versioned archive. OpenKIM exists precisely to make potentials citable and versioned, and its use removes a class of ambiguity.
The general principle: anything that affects the numerical result and is not the engine binary belongs in the repository as a committed file.
Reproducibility Beyond the Repository: Seeds, Hardware, and Floating Point
A molecular dynamics repository can be perfectly organized but still fail to reproduce a result, because molecular dynamics has sources of non-determinism that live outside of version control. Understanding them avoids false confidence.
Random seeds control the initial velocity assignment and, in stochastic methods, the behavior of the thermostat and barostat. If the seed is not saved, the run is not reproducible even with identical inputs. Many engines accept an explicit seed; use it and save it.
Parallel decomposition affects the order of floating point summation. Running the same system on 16 cores versus 64 cores can produce trajectories that diverge over time due to accumulated rounding differences. This is not a bug; this is the nature of floating point arithmetic under different orders of reduction. For strict reproducibility, record the domain decomposition and number of cores, or accept that bitwise identity is not feasible and aim for statistical reproducibility instead.
Hardware and compiler differences introduce further variation through different math libraries and instruction sets. Containers reduce but do not eliminate this.
Thermostat and barostat choices change the ensemble and therefore the physics. A repository should record the ensemble explicitly — NVT, NPT, NVE — along with the coupling constants.
The honest framework is that bitwise reproducibility is achievable in a fixed hardware and software configuration, and that statistical reproducibility is the realistic goal across configurations. A repository that documents the configuration makes the first achievable and the second verifiable.
FAIR Principles Applied to Simulation Data
The FAIR principles – Findable, Accessible, Interoperable, Reusable – were formulated for research data in general and translate into specific molecular dynamics repository practices.
Findable means the dataset has a persistent identifier and descriptive metadata. A DOI from Zenodo or an institutional repository satisfies this; this is not the case for a directory on a laboratory server.
Accessible means the data can be retrieved by a human or machine using a standard protocol, with a clear license. Choosing an open license at the time of deposit avoids the common situation where data is archived but legally unusable.
Interoperable means that the formats are standard and documented. Using XTC, DCD or NetCDF rather than a custom binary format and documenting the units makes the trajectories readable by standard tools years later.
Reusable means there is sufficient context to reuse the data for a new purpose. This requires the metadata described above: engine version, force field version, ensemble, temperature, and any applied processing.
The FAIR framing is useful because it shifts attention from “did I save the files” to “can someone else use these files.” Those are different questions, and only the second one matters for the long data.
Practical Setup: A Minimal Working Example
A minimal reproducible setup for a GROMACS-style molecular dynamics repository workflow, adaptable to other engines, looks like this in practice.
Initialize the repository and configure large-file handling:
git init md-project
cd md-project
git lfs install
git lfs track "*.xtc" "*.trr" "*.tpr"
git add .gitattributes
Create the environment specification and commit it alongside the inputs:
conda env export --no-builds > environment.yml
git add environment.yml systems/ simulations/ analysis/
git commit -m "Initial reproducible setup"
Label the versions so that a DOI can be created and archive the frozen dataset separately from the working repository. The tag marks the exact state of the code; the archive contains the exact state of the data.
The workflow above is deliberately minimal. The point is not the sophistication of the tools, but the discipline of committing the elements that determine the outcome and archiving the elements that are too large to version.
Sources & Further Reading
- Molecular dynamics — Wikipedia: Molecular dynamics (MD) is a computer simulation method for analyzing the physical movements of atoms and molecules. The atoms and molecules are allowed to interact…
Frequently Asked Questions
What is a molecular dynamics repository?
A molecular dynamics repository is a version-controlled store for the files that define and reproduce a simulation: engine inputs, topology and coordinate files, force-field parameters, run scripts, and analysis code. The term also refers to the upstream source repositories of engines like GROMACS and LAMMPS, and to data archives holding trajectories. Distinguishing these three uses prevents most repository-design mistakes.
Should I store MD trajectories in Git?
Trajectories generally should not go into plain Git history, because Git stores every version permanently and large binaries bloat the repository irreversibly. Git LFS works for moderately sized files that change rarely, and DVC works better for large data that is regenerated frequently. The published, frozen version of a trajectory belongs in a DOI-issuing data repository such as Zenodo.
How do I make a molecular dynamics simulation reproducible?
Reproducibility requires pinning the engine version, force field version, random seed, exact input files, and hardware or container configuration. Validating force field parameter files in the repository removes the most common source of silent divergence. Bitwise reproducibility is realistic in a fixed configuration; on different numbers of cores or hardware, aim for statistical reproducibility and document the difference.
What trajectory format should I use?
XTC is a good default for GROMACS workflows because it compresses well with configurable precision. DCD is widely readable and uncompressed, which is suitable for interoperability. TRR preserves full precision and is useful when speeds or forces are large. NetCDF-based formats are self-describing and contain unit metadata, which facilitates long-term reuse.
Do I need a DOI for my simulation data?
A DOI is necessary if you want the dataset to be citable independently of the paper, which is increasingly expected by journals and funders. Zenodo and Figshare both mint DOIs and integrate with GitHub releases, so tagging a release can produce a citable identifier automatically. For internal, unpublished work, a DOI is optional but the same archiving discipline still pays off.
How should force-field versions be recorded?
Force field parameter files must be sold in the repository under a parameter directory and validated, because they are small text files and their exact content determines the result. The force field name and revision must also be recorded in the README and in a machine-readable metadata file. Any local changes to fees or related settings should be validated rather than kept in a home directory.
P.S. A few readers have asked which university-backed specializations we actually reach for — it's Coursera Plus; if you want the current details.
Frequently asked questions
What is a molecular dynamics repository?
A molecular dynamics repository is a version-controlled store for the files that define and reproduce a simulation: engine inputs, topology and coordinate files, force-field parameters, run scripts, and analysis code. The term also refers to the upstream source repositories of engines like GROMACS and LAMMPS, and to data archives holding trajectories. Distinguishing these three uses prevents most repository-design mistakes.
Should I store MD trajectories in Git?
Trajectories generally should not go into plain Git history, because Git stores every version permanently and large binaries bloat the repository irreversibly. Git LFS works for moderately sized files that change rarely, and DVC works better for large data that is regenerated frequently. The published, frozen version of a trajectory belongs in a DOI-issuing data repository such as Zenodo.
How do I make a molecular dynamics simulation reproducible?
Reproducibility requires pinning the engine version, force field version, random seed, exact input files, and hardware or container configuration. Validating force field parameter files in the repository removes the most common source of silent divergence. Bitwise reproducibility is realistic in a fixed configuration; on different numbers of cores or hardware, aim for statistical reproducibility and document the difference.
What trajectory format should I use?
XTC is a good default for GROMACS workflows because it compresses well with configurable precision. DCD is widely readable and uncompressed, which is suitable for interoperability. TRR preserves full precision and is useful when speeds or forces are large. NetCDF-based formats are self-describing and contain unit metadata, which facilitates long-term reuse.
Do I need a DOI for my simulation data?
A DOI is necessary if you want the dataset to be citable independently of the paper, which is increasingly expected by journals and funders. Zenodo and Figshare both mint DOIs and integrate with GitHub releases, so tagging a release can produce a citable identifier automatically. For internal, unpublished work, a DOI is optional but the same archiving discipline still pays off.
How should force-field versions be recorded?
Force field parameter files must be sold in the repository under a parameter directory and validated, because they are small text files and their exact content determines the result. The force field name and revision must also be recorded in the README and in a machine-readable metadata file. Any local changes to fees or related settings should be validated rather than kept in a home directory.
Earn certificates from real universities
One subscription for university-backed Python and data-science certificates