What is research software engineering?
Research software engineering (RSE) is the discipline of applying professional software engineering practices — version control, testing, code review, continuous integration, documentation, and long-term maintenance — to the code that underpins scientific research. If you have ever spent three weeks debugging a molecular dynamics analysis script that only runs on one postdoc’s laptop, or tried to reproduce a published result from a tarball of undocumented Python, you have already met the problem that research software engineering exists to solve. This article explains what RSE is, how it differs from adjacent roles, what a research software engineer actually does day to day, and how to decide whether your group needs one.
The short definition, and why it is contested
The term “research software engineer” was coined in the UK around 2012, largely through the work of the Software Sustainability Institute, to give a name to a growing population of people who were neither traditional researchers nor IT support staff. They wrote code that was essential to research, but their career progression, job titles, and recognition did not fit either box.
The commonly cited working definition is that a research software engineer is someone who combines:
- A professional understanding of software engineering — design, testing, version control, deployment, maintenance.
- An active role in research — contributing to research questions, publications, and grant applications, not merely implementing a spec handed down by someone else.
- Deep engagement with at least one research domain — enough to understand the science, the data, and the constraints.
The contested part is the boundary. Is a bioinformatician who maintains a widely used R package an RSE? Is a physicist who writes simulation code for their own papers an RSE?
Is a research group’s “data manager” who writes a lot of Python an RSE? There is no licensing body and no universally agreed certification, so the label is applied pragmatically. In practice, the distinguishing feature is orientation: RSEs treat software as a first-class research output and a long-lived artifact, rather than as a disposable means to a single paper.
RSE vs. adjacent roles: a comparison
The fastest way to understand RSE is to see where it sits relative to roles it is often confused with. The table below is a heuristic, not a taxonomy — real people blur these lines constantly.
Related: — Project-based data-science paths with a guided terminal and real datasets.
| Dimension | Research Software Engineer | Traditional Researcher (who codes) | IT / Research Computing Support | Data Scientist / Analyst |
|---|---|---|---|---|
| Primary output | Maintainable, reusable software | Publications and findings | Working infrastructure and services | Models, analyses, decisions |
| Code lifespan | Years to decades; versioned releases | Often tied to one project/paper | Service lifetime; often vendor-managed | Often project-bound |
| Typical skills | Software design, testing, CI/CD, HPC, domain science | Domain science, methods, statistics | Systems admin, networking, schedulers, security | Statistics, ML, domain data |
| Relationship to research question | Co-defines and shapes it | Owns it | Serves it indirectly | Often answers a posed question |
| Success metric | Adoption, reproducibility, citations of software, sustainability | Publications, grants, impact | Uptime, user satisfaction, cost | Business/research insight |
| Career track | Often hybrid; sometimes permanent staff scientist | Faculty / postdoc ladder | IT ladder | Industry or academic |
The key insight: an RSE is not “a researcher who happens to be good at code,” nor “an IT person who helps with Python.” The role is defined by owning the software as a research product while remaining embedded in the science.
What a research software engineer actually does
Job descriptions vary enormously, but the work clusters into recognizable categories. A realistic week might touch four or five of these.
1. Building and maintaining research software
This is the core. It includes designing APIs for simulation codes, refactoring a monolithic analysis script into a tested library, packaging tools for distribution on PyPI or conda-forge, and keeping dependencies from breaking. For a molecular dynamics group, this might mean maintaining a GROMACS or LAMMPS analysis toolkit, or writing a NumPy/HDF5 pipeline that streams trajectory data without loading terabytes into RAM.
Worth a look: — One subscription for university-backed Python and data-science certificates.
2. Making research reproducible
RSEs are often the people who introduce environment.yml, pyproject.toml, container images (Docker, Singularity/Apptainer), and workflow managers (Snakemake, Nextflow, Common Workflow Language) into a lab. The goal is that a result can be regenerated on a different machine, by a different person, two years later. This is where RSE overlaps heavily with the reproducibility movement and with FAIR data principles.
3. Performance and scaling
Scientific code frequently needs to run faster or bigger. RSEs profile code, vectorize NumPy operations, parallelize with MPI or OpenMP, port hot loops to C/C++/Fortran or use Numba/Cython, and tune I/O — HDF5 chunking and compression choices, for example, can dominate runtime in trajectory analysis. They also help groups use HPC clusters and GPUs effectively.
4. Consulting and training
Many RSEs run office hours, code reviews, and workshops (Software Carpentry, HPC Carpentry, domain-specific training). A large fraction of their impact comes from teaching researchers to write better code themselves, rather than doing it all for them.
5. Grant and publication contributions
RSEs increasingly appear as co-authors and as named personnel on grants, because funders now expect a software sustainability and management plan. Writing that plan, estimating effort, and committing to maintenance are RSE activities.
Where RSEs sit organizationally
There are three common models, each with trade-offs:
- Centralized RSE group (a university-wide or institute-wide team). Advantages: shared expertise, career path for RSEs, ability to staff multiple projects. Disadvantages: can be distant from any single research question; chargeback or prioritization politics.
- Embedded in a research group or lab. Advantages: deep domain context, fast iteration, strong relationships. Disadvantages: isolation, no peer RSE community, risk that the RSE becomes a general “fix my computer” resource.
- Hybrid / matrix. A central group with staff seconded to projects. Increasingly common; requires clear line management and time-allocation agreements.
The UK has a particularly visible RSE community, with an annual RSE Conference and regional networks. Similar communities exist in Germany (de-RSE), the Netherlands, the US, and Australia. The Society of Research Software Engineering (Society of RSE) was established in the UK as a professional body.
How RSE relates to reproducibility and open science
The reproducibility crisis in computational science is largely a software problem. Analyses depend on specific library versions, undocumented preprocessing steps, and code that was never released. RSE practices directly address this:
- Version control and tagged releases make it possible to cite the exact software version used.
- Testing catches the silent numerical bugs that produce plausible-but-wrong results.
- Continuous integration ensures code still works after dependency updates.
- Archiving code in Zenodo or Software Heritage gives it a DOI and a permanent home.
- Documentation lets others (including future you) understand and reuse it.
The FAIR principles for research software, developed by the Research Data Alliance, extend the FAIR data guidelines explicitly to software. Journals increasingly require code availability statements, and some require code review as part of publication.
How to decide whether you need an RSE
Not every group needs a dedicated RSE. Use these criteria as a rough filter:
You probably need one if:
- Software is central to your research output and is used by people beyond your group.
- You maintain code across multiple grants and years, and it keeps breaking.
- You are repeatedly rebuilding the same infrastructure (I/O, parallelism, packaging) that others have already solved.
- Reviewers, collaborators, or funders ask about software sustainability and you have no good answer.
- Your group spends significant time on software maintenance that is invisible in publications.
You may not need one (yet) if:
- Your code is genuinely throwaway — a one-off analysis for a single figure.
- Your needs are met by well-maintained community tools and you only write glue scripts.
- You have a group member who genuinely enjoys and is good at this, and has time.
A middle ground: hire a part-time CSR, participate in a core group, or invest in training to help existing researchers adopt best practices. The worst outcome would be to hire a CSR and then use them as ad hoc IT support.
Career paths and how to become an RSE
CSR careers are truly diverse. Entry points include:
- PhD in a computational domain (physics, chemistry, bioinformatics) who gravitated toward the software side.
- Software engineer from industry who wants to work on scientific problems.
- Research staff who formalize their software role.
Skills that matter most, roughly in order of how often they come up:
- Knowledge of Python (and often C++, Fortran, or Julia for performance-critical code).
- Version control with Git, including branching and code review workflows.
- Test frameworks (Pytest, GoogleTest) and CI (GitHub Actions, GitLab CI).
- Packaging and environmental management (Conda, Pip, Container).
- HPC and parallel programming (MPI, OpenMP, CUDA, job schedulers like Slurm).
- Scientific competence in the field: sufficient to speak with researchers as colleagues.
- Communication and teaching skills.
The career ladder is less standardized than in academia or industry. Some RSEs become permanent staff scientists; some move into research computing leadership; some return to faculty; some go to industry. The lack of a clear ladder is a real, frequently discussed problem in the community, and one reason professional bodies and central groups matter.
Common misconceptions
- “RSE is just IT support for scientists.” No — RSEs shape research questions and own software products. IT support maintains infrastructure.
- “Any good programmer can do RSE.” Domain context matters enormously. A brilliant web developer may not know why a floating-point summation order changes a result.
- “RSEs only write code.” A large share of the job is consulting, training, and design discussion.
- “RSE is a stepping stone to a ‘real’ academic job.” For some it is; for many it is a legitimate, permanent career.
- “You need a computer science degree.” Most RSEs come from scientific domains, not CS.
Key Takeaways
- Research software engineering applies professional software practices to research code, treating software as a first-class, long-lived research output.
- An RSE is defined by combining software engineering skill, active research involvement, and domain expertise — not by a specific job title or certification.
- RSE differs from IT support, data science, and “researchers who code” primarily in ownership of software as a product and its multi-year lifespan.
- Core activities include building and maintaining code, enabling reproducibility, performance tuning, consulting, training, and contributing to grants and publications.
- RSEs can be centralized, embedded, or hybrid; each model has real trade-offs in context, career path, and prioritization.
- Deciding whether you need an RSE depends on whether software is central, reused, and long-lived — not on group size alone.
Sources & Further Reading
- Research software engineering — Wikipedia: Research software engineering is the application of software engineering practices, methods and techniques for research software, i.e. software that was made for…
- Software engineering — Wikipedia: Software engineering is a branch of both computer science and engineering focused on designing, developing, testing, and maintaining software applications. It involves…
Frequently Asked Questions
What is research software engineering in simple terms?
Research software engineering is the practice of building and maintaining the software that scientists rely on, using the same professional standards — testing, version control, documentation, maintenance — that commercial software teams use. A research software engineer works at the intersection of software development and a scientific domain, so they understand both the code and the science it serves. The goal is software that is correct, reusable, and still works years later.
How is a research software engineer different from a software engineer?
A traditional software developer typically creates products for users or customers, and requirements are often determined by a product team. A research software developer works on software whose purpose is to generate or support scientific knowledge, where the requirements are often exploratory in nature and the “correct” result may be unknown in advance. RSEs also typically require actual domain knowledge, enough to judge whether a numerical result is physically plausible, not just whether the code will execute.
Do I need a PhD to become a research software engineer?
No, but domain knowledge is essential, and a PhD is one common way to acquire it. Many RSEs hold PhDs in physics, chemistry, bioinformatics, or related fields and moved toward the software side of their work. Others come from industry software engineering and learn the domain on the job. What matters is being able to engage with researchers as a peer about the science, not the credential itself.
What programming languages do research software engineers use?
Python dominates, especially for analysis, scripting, and glue code, often alongside NumPy, pandas, and HDF5-based I/O. Performance-critical components are frequently written in C++, Fortran, or C, with Julia and Rust gaining ground. R remains important in statistics and bioinformatics, and shell scripting, Make, and workflow languages like Snakemake and Nextflow are ubiquitous. The language matters less than the practices around it.
Is research software engineering a good career?
For people who enjoy both science and software, it can be an excellent career with strong demand, varied problems, and visible impact. The main caveat is that career structures are less standardized than in academia or industry, so progression depends heavily on the institution and whether a central RSE group exists. Many RSEs find the work more stable and better paid than postdoc positions, though the lack of a universal ladder is a recognized issue.
How does research software engineering support reproducibility?
RSE practices make results reproducible by pinning exact software versions, automating builds and tests, documenting dependencies, and archiving code with a DOI so it can be cited and retrieved. Without these practices, analyses often depend on undocumented steps and specific library versions that are impossible to reconstruct later. RSEs also introduce containers and workflow managers that capture an entire computational environment, not just the code.
Further reading and authoritative sources
- The Wikipedia article on research software engineering provides a broad overview and history of the term.
- The Software Sustainability Institute (software.ac.uk) publishes extensively on RSE definitions, careers, and community surveys.
- The Society of Research Software Engineering (society-rse.org) is the UK professional body for the field.
- The FAIR Principles for Research Software, from the Research Data Alliance, extend FAIR data guidelines to software specifically.
- The Journal of Open Source Software (joss.theoj.org) and the Journal of Open Research Software publish peer-reviewed research software.
P.S. A few readers have asked which interactive course platform we actually reach for — it's DataCamp; if you want the current details.
Frequently asked questions
What is research software engineering in simple terms?
Research software engineering is the practice of building and maintaining the software that scientists rely on, using the same professional standards — testing, version control, documentation, maintenance — that commercial software teams use. A research software engineer works at the intersection of software development and a scientific domain, so they understand both the code and the science it serves. The goal is software that is correct, reusable, and still works years later.
How is a research software engineer different from a software engineer?
A traditional software developer typically creates products for users or customers, and requirements are often determined by a product team. A research software developer works on software whose purpose is to generate or support scientific knowledge, where the requirements are often exploratory in nature and the 'correct' result may be unknown in advance. RSEs also typically require actual domain knowledge, enough to judge whether a numerical result is physically plausible, not just whether the code will execute.
Do I need a PhD to become a research software engineer?
No, but domain knowledge is essential, and a PhD is one common way to acquire it. Many RSEs hold PhDs in physics, chemistry, bioinformatics, or related fields and moved toward the software side of their work. Others come from industry software engineering and learn the domain on the job. What matters is being able to engage with researchers as a peer about the science, not the credential itself.
What programming languages do research software engineers use?
Python dominates, especially for analysis, scripting, and glue code, often alongside NumPy, pandas, and HDF5-based I/O. Performance-critical components are frequently written in C++, Fortran, or C, with Julia and Rust gaining ground. R remains important in statistics and bioinformatics, and shell scripting, Make, and workflow languages like Snakemake and Nextflow are ubiquitous. The language matters less than the practices around it.
Is research software engineering a good career?
For people who enjoy both science and software, it can be an excellent career with strong demand, varied problems, and visible impact. The main caveat is that career structures are less standardized than in academia or industry, so progression depends heavily on the institution and whether a central RSE group exists. Many RSEs find the work more stable and better paid than postdoc positions, though the lack of a universal ladder is a recognized issue.
How does research software engineering support reproducibility?
RSE practices make results reproducible by pinning exact software versions, automating builds and tests, documenting dependencies, and archiving code with a DOI so it can be cited and retrieved. Without these practices, analyses often depend on undocumented steps and specific library versions that are impossible to reconstruct later. RSEs also introduce containers and workflow managers that capture an entire computational environment, not just the code. Further reading and authoritative sources - The Wikipedia article on research software engineering provides a broad overview and history of th
Learn Python by coding in your browser
Interactive Python and data-science courses you code directly in the browser