Skip to main content Skip to secondary navigation

Data Storage

Main content start

For Active Research Data

Storage on Sherlock

The Sherlock compute cluster includes a file system for storing data. Each user and Principal Investigator (PI) group is given access to the pre-defined directories $SCRATCH and $GROUP_SCRATCH for storing data and code directly associated with projects for which you are actively using Sherlock’s computational resources. The maximum retention period (for data that remains unchanged) is 90 days. Because scratch storage on Sherlock is intended for active compute rather than as a backup target, longer-lived data is better placed on Oak.

Learn more: Storage on Sherlock

Oak Storage Service

Seamlessly integrated with the Sherlock and SCG compute clusters, Oak offers PIs and their researchers an affordable, scalable persistent storage tier for active research data. It is well suited to large datasets that are periodically staged to scratch storage for active compute, curated post-processed results from job campaigns, and final outputs underlying publications. Oak is designed for data with a long working lifetime that nonetheless needs to remain readily accessible, distinguishing it from purely cold archival storage. 

Rather than relying on a single vendor’s solution, Oak is built on open-source technologies (Lustre and the Robinhood Policy Engine) and was designed in-house by the Stanford Research Computing team. Oak supports a wide range of access methods for moving and managing data, including Globus, SFTP, RSYNC, SCP, SSHFS, and RCLONE, with SMB and NFSv4 gateways available for a small additional fee. 

Oak is approved to store Low and Moderate Risk Data as defined by the Information Security Office, and is billed monthly. 

Learn more: Oak Storage Service | Rates | Oak Gateways | Oak Technical Documentation 

For Long-term Data Archiving

Elm Cold Storage Service

Elm is designed for long-term cold storage of large research datasets, from terabytes to hundreds of petabytes. It addresses the need for inexpensive, dependable archival storage to preserve research data for future use or to meet regulatory and policy compliance obligations. Unlike Oak, which keeps active datasets readily accessible for compute, Elm is optimized for data that is written once and retrieved infrequently, if ever.

Data is safely and securely stored on a state-of-the-art tape system in our data center, with high-throughput Globus transfers available for moving data in and out. Elm offers significant savings over cloud-based archival storage, with no egress or restore fees beyond the monthly allocation cost.

Elm is currently approved to store Low and Moderate Risk Data as defined by the Information Security Office. Efforts to provide an Elm for High Risk / PHI service are underway; for updates, email srcc-support@stanford.edu.

Learn more: Elm Cold Storage Service | Rates | Elm Technical Documentation

Oak or Elm? Help me choose ...

Oak — Active Research

Elm — Cold Archive

Primary intentHigh-performance storage for data you use and modify daily.Long-term retention of completed projects or raw data.
WorkflowSeamlessly integrated with Sherlock and SCG. Ideal for large datasets that are periodically staged to scratch storage for active compute.Not integrated with Sherlock and SCG. Data is moved once, manually, and rarely touched again.
Best for ...Running simulations, data analysis, and active code development.Storing “finished” datasets to meet grant or publication requirements.
Retrieval speedInstant. Built for high-speed computation.Slow. Tape-based system; large restores can take hours or days.
Scheduled backups?No. While high-performance, it does not provide automated snapshots or versioning.No. Not for automated/scheduled workflows. Tape drives cannot handle frequent, small writes.
Crucial limitIf you delete a file, it is gone.Strictly for infrequently accessed data (less than once or twice per year).

Next steps →

Oak Storage Service

Elm Cold Storage Service