[Dataset name]
90 days of Windows authentication events from a 400-host lab network, with 14 labeled lateral-movement sequences.
12.4 GB · Parquet · CC-BY-4.0 · Updated [date]
Seacurity hosts large security datasets, the software to analyze them, and the compute to run that software. Download everything or work with it in place. There is no paid version.
Browsing and downloads don't require an account.
Useful datasets are sold under vendor contracts. Useful tooling assumes you already have a cluster and someone to run it. If you have neither, you learn on small synthetic samples and hope the lessons transfer to production.
We think the data and the tools should be public, so that's what Seacurity is.
Logs, packet captures, threat intelligence feeds, and recorded attack traces, at the volumes you'd see in a real environment. Each dataset has a schema, a version history, and a license. Download it or query it where it sits.
Open-source detection engines, parsers, hunting notebooks, and visualizers, installed and configured. The same code is on GitHub if you'd rather run it yourself.
Hosted workspaces with the datasets already mounted. You can run a query over a few billion events without provisioning anything first.
1dataset: win-auth-lab/v32source: instrumented lab network3hosts: 4004span: 90 days5labels: 14 lateral-movement sequences6license: CC-BY-4.07covers: 4624, 4625, 4648, 46728excludes: Kerberos pre-auth failures
The people with the least access to real security data are the ones who could use it most. The data exists. Most of it is under contract.
Reproducible, citable data for papers and benchmarks.
Every dataset here has a fixed version you can reference.
90 days of Windows authentication events from a 400-host lab network, with 14 labeled lateral-movement sequences.
12.4 GB · Parquet · CC-BY-4.0 · Updated [date]
[One line stating exactly what it contains.]
[Size] · [Format] · [License] · Updated [date]
[One line stating exactly what it contains.]
[Size] · [Format] · [License] · Updated [date]
Detection engines, parsers, hunting notebooks, and visualizers are ready in every workspace. The same code is on GitHub if you'd rather run it yourself.
Nothing is paid. There are no usage tiers or trial periods, and nothing on the site requires talking to a salesperson.
Datasets carry a license. Tools are open source. The collection and cleaning methods are written up so you can reproduce them.
Large, noisy, inconsistently formatted. Cleaning that up is part of the work, and you should get to practice it.
Every dataset is reviewed. Where we can't anonymize confidently, we generate synthetic data or leave the field out.
Roadmap decisions are made in the open, and contributions are credited by name.
If you have a dataset that could be released, a tool worth adding, or a tutorial to write, the contributor guide explains how to submit it and what review looks like.
A dataset that could be released, with a data card.
Open-source software worth installing for everyone.
A walkthrough that teaches with real data.
Every submission is reviewed and credited by name.
# datacard.yamlname: your-datasetsource: contributed | lab | syntheticcollection: how it was gatheredcoverage: what it includesexcludes: what it does notlicense: CC-BY-4.0pii_review: required before publish
The people with the least access to real security data are the ones who could use it most.
Graduate students, small security teams, independent researchers, people teaching the subject. The data exists. Most of it is under contract.
Seacurity is an attempt to fix that by collecting what can be released, cleaning it, and hosting it with the tools needed to use it.
[Team / organization]
With support from [partners]. The code and data are licensed so that the project can continue without us.
Partners
Nothing here requires talking to a salesperson. If it isn't covered, ask in the discussion forum.
Yes. Seacurity is funded by [grants / sponsors / donations]. There is no paid tier and no plan to add one.
Depends on the dataset. The license is on each data card. Most are permissive; some are research-only.
Three sources: corpora contributed by organizations, instrumented lab environments we run, and synthetic generation. The data card for each dataset says which.
It shouldn't be. Every release is reviewed and personal data is anonymized or removed. If you find something we missed, report it and we'll pull the affected version.
Not to browse or download. Workspaces and contributions require one so we can save your work and attribute it.
Email us with a description of the project. We extend resources for research groups, courses, and open-source work on a case-by-case basis.
Yes. Everything is on GitHub with packaging for self-hosting. The hosted version exists to skip setup, not to lock you in.
One email when we publish new datasets or tools. Roughly monthly.
No account needed. Unsubscribe in one click.