Open data. Open water.

Security data and tools,free for anyone who wants them.

Seacurity hosts large security datasets, the software to analyze them, and the compute to run that software. Download everything or work with it in place. There is no paid version.

Browsing and downloads don't require an account.

The problem

Most security research starts with a procurement form.

Useful datasets are sold under vendor contracts. Useful tooling assumes you already have a cluster and someone to run it. If you have neither, you learn on small synthetic samples and hope the lessons transfer to production.

We think the data and the tools should be public, so that's what Seacurity is.

What's here
01

Data

Logs, packet captures, threat intelligence feeds, and recorded attack traces, at the volumes you'd see in a real environment. Each dataset has a schema, a version history, and a license. Download it or query it where it sits.

02

Tools

Open-source detection engines, parsers, hunting notebooks, and visualizers, installed and configured. The same code is on GitHub if you'd rather run it yourself.

AB
03

Compute

Hosted workspaces with the datasets already mounted. You can run a query over a few billion events without provisioning anything first.

How it works

Pick a dataset.
Open a workspace. Work.

data-card.yaml
1dataset: win-auth-lab/v3
2source: instrumented lab network
3hosts: 400
4span: 90 days
5labels: 14 lateral-movement sequences
6license: CC-BY-4.0
7covers: 4624, 4625, 4648, 4672
8excludes: Kerberos pre-auth failures
Reviewed for personal data
Who it's for

Who uses
Seacurity.

The people with the least access to real security data are the ones who could use it most. The data exists. Most of it is under contract.

Reproducible, citable data for papers and benchmarks.

Every dataset here has a fixed version you can reference.

Audience01 / 05
Catalog

Featured
datasets.

Full catalog
01Reviewed for personal data

[Dataset name]

90 days of Windows authentication events from a 400-host lab network, with 14 labeled lateral-movement sequences.

12.4 GB · Parquet · CC-BY-4.0 · Updated [date]

02Reviewed for personal data

[Dataset name]

[One line stating exactly what it contains.]

[Size] · [Format] · [License] · Updated [date]

03Reviewed for personal data

[Dataset name]

[One line stating exactly what it contains.]

[Size] · [Format] · [License] · Updated [date]

Tools

Installed, configured,
and open source.

Detection engines, parsers, hunting notebooks, and visualizers are ready in every workspace. The same code is on GitHub if you'd rather run it yourself.

Elasticsearch
Search & storage
Kibana
Visualizer
Sigma
Detection rules
Zeek
Network analysis
Suricata
Detection engine
Jupyter
Hunting notebooks
Spark
Query engine
Sysmon
Endpoint telemetry
YARA
Pattern matching
Wireshark
Packet analysis
MISP
Threat intelligence
OpenSearch
Search & storage
Elasticsearch
Search & storage
Kibana
Visualizer
Sigma
Detection rules
Zeek
Network analysis
Suricata
Detection engine
Jupyter
Hunting notebooks
Spark
Query engine
Sysmon
Endpoint telemetry
YARA
Pattern matching
Wireshark
Packet analysis
MISP
Threat intelligence
OpenSearch
Search & storage
OpenSearch
Search & storage
MISP
Threat intelligence
Wireshark
Packet analysis
YARA
Pattern matching
Sysmon
Endpoint telemetry
Spark
Query engine
Jupyter
Hunting notebooks
Suricata
Detection engine
Zeek
Network analysis
Sigma
Detection rules
Kibana
Visualizer
Elasticsearch
Search & storage
OpenSearch
Search & storage
MISP
Threat intelligence
Wireshark
Packet analysis
YARA
Pattern matching
Sysmon
Endpoint telemetry
Spark
Query engine
Jupyter
Hunting notebooks
Suricata
Detection engine
Zeek
Network analysis
Sigma
Detection rules
Kibana
Visualizer
Elasticsearch
Search & storage
Principles

How we
operate.

Nothing is paid. There are no usage tiers or trial periods, and nothing on the site requires talking to a salesperson.

No paid tierNo trialsNo sales callsOpen sourceLicensed data

Everything is documented

Datasets carry a license. Tools are open source. The collection and cleaning methods are written up so you can reproduce them.

We publish data the way it actually looks

Large, noisy, inconsistently formatted. Cleaning that up is part of the work, and you should get to practice it.

Personal data is removed before release

Every dataset is reviewed. Where we can't anonymize confidently, we generate synthetic data or leave the field out.

Contributors set direction

Roadmap decisions are made in the open, and contributions are credited by name.

Contribute

Seacurity runs on
contributed data and code.

If you have a dataset that could be released, a tool worth adding, or a tutorial to write, the contributor guide explains how to submit it and what review looks like.

Data

A dataset that could be released, with a data card.

Tools

Open-source software worth installing for everyone.

Tutorials

A walkthrough that teaches with real data.

Review

Every submission is reviewed and credited by name.

# datacard.yaml
name: your-dataset
source: contributed | lab | synthetic
collection: how it was gathered
coverage: what it includes
excludes: what it does not
license: CC-BY-4.0
pii_review: required before publish
About

The people with the least access to real security data are the ones who could use it most.

Graduate students, small security teams, independent researchers, people teaching the subject. The data exists. Most of it is under contract.

Seacurity is an attempt to fix that by collecting what can be released, cleaning it, and hosting it with the tools needed to use it.

Maintained by

[Team / organization]

With support from [partners]. The code and data are licensed so that the project can continue without us.

Partners

[Partner][Partner][Partner][Partner][Partner][Partner]
[Partner][Partner][Partner][Partner][Partner][Partner]
FAQ

Questions,
answered.

Nothing here requires talking to a salesperson. If it isn't covered, ask in the discussion forum.

Yes. Seacurity is funded by [grants / sponsors / donations]. There is no paid tier and no plan to add one.

Depends on the dataset. The license is on each data card. Most are permissive; some are research-only.

Three sources: corpora contributed by organizations, instrumented lab environments we run, and synthetic generation. The data card for each dataset says which.

It shouldn't be. Every release is reviewed and personal data is anonymized or removed. If you find something we missed, report it and we'll pull the affected version.

Not to browse or download. Workspaces and contributions require one so we can save your work and attribute it.

Email us with a description of the project. We extend resources for research groups, courses, and open-source work on a case-by-case basis.

Yes. Everything is on GitHub with packaging for self-hosting. The hosted version exists to skip setup, not to lock you in.

Newsletter

Release
notes.

One email when we publish new datasets or tools. Roughly monthly.

No account needed. Unsubscribe in one click.