Schema Markup Guide

Dataset Schema: Complete Reference

Schema MarkupPublished Jul 17, 2026Updated Jul 19, 20263 min readLinkedInX

Dataset schema describes a collection of data — a research dataset, a public data table, an API-accessible resource — making it discoverable in Google Dataset Search, a specialized engine for finding data. It is a niche type with a clear audience: researchers, governments, and organizations that publish original data. For everyone else it is rarely relevant. This reference covers the required fields, the access and license properties, and who genuinely benefits.

What Dataset Schema Does

It makes a dataset findable in Google Dataset Search and helps Google understand the data’s subject, provenance, and access terms — turning a data page into a discoverable research resource.

Dataset Search is a distinct product from web search, used by people specifically looking for data. If you publish original datasets — as this site would for its own research studies — Dataset markup is how those datasets get surfaced there.

Required and Recommended Fields

The essentials are name and description; strongly recommended are creator, license, distribution (how to access the data), temporalCoverage, spatialCoverage, and variableMeasured.

A clear, specific description is what makes a dataset discoverable and understandable. distribution points to the actual downloadable file or access endpoint with a format. The coverage and variable properties help users judge relevance. Validate with the Rich Results Test or the Dataset-specific guidance.

Access and License Properties

Use license to state usage terms, distribution (a DataDownload with contentUrl and encodingFormat) to declare how to get the data, and isAccessibleForFree to indicate whether it is free.

These properties tell both Google and researchers how the data can be used and obtained — critical for a resource whose whole value is being reusable. Point license at real terms and distribution at a genuine file or endpoint.

Who Actually Benefits

Researchers, academic institutions, governments, and organizations publishing original data benefit; typical marketing or content sites do not, because they have no datasets to expose.

Do not add Dataset markup to ordinary pages — it applies only to genuine data resources. Where it fits, it is one of the safe, additive types. For a site publishing original research, it is the natural complement to the study methodology in the Research & Data pillar.

Example Markup

A minimal Dataset (shown as code):

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Dataset",
  "name": "Example AI Citation Study Data",
  "description": "Monthly counts of brand citations across major AI assistants.",
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "creator": { "@type": "Organization", "name": "Example Co" },
  "distribution": {
    "@type": "DataDownload",
    "encodingFormat": "text/csv",
    "contentUrl": "https://example.com/data.csv"
  }
}
</script>

See JSON-LD vs Microdata for the format.

Key Takeaways

  • Dataset schema makes data discoverable in Google Dataset Search, a separate engine for finding data.
  • Required: name, description. Recommended: creator, license, distribution, temporal/spatial coverage, variableMeasured.
  • Use distribution (DataDownload with contentUrl + encodingFormat) and license to declare access and terms.
  • It benefits researchers, institutions, governments, and original-data publishers — not typical content sites.
  • Only add it to genuine data resources; a clear description drives discoverability.

Frequently Asked Questions

What is Google Dataset Search?

A specialized search engine, separate from web search, for finding datasets. Dataset schema is what makes your data appear in it — useful if you publish original research or public data, and irrelevant if you do not.

Do I need Dataset schema for my blog?

Almost certainly not. Dataset markup applies to genuine data resources — research datasets, data tables, downloadable files. A typical blog or marketing site has no datasets to expose and should not add it.

What is the distribution property?

It describes how to actually access the data — a DataDownload with a contentUrl (the file location) and encodingFormat (like text/csv). It is what turns a described dataset into a genuinely obtainable one for users and Google.

The Bottom Line

Dataset schema is a specialist type for a specialist audience: if you publish original data, it makes that data discoverable and reusable in Dataset Search; if you do not, skip it. Add it to real datasets with clear descriptions and honest access terms — the same fit-to-purpose judgment the whole schema cluster applies to every type.


Further reading & sources

See how your site actually shows up in AI search. An AI visibility audit maps where you’re cited, where you’re invisible, and what to fix first — in plain English.

Get your AI visibility auditTry the free SEO tools →

Prefer self-serve? The interactive checklists turn guides like this one into a working to-do list.

Keep reading in Schema Markup

Get one email when something genuinely changes

AI search moves fast and most of it is noise. We send one short email when a real shift is worth your time. Unsubscribe anytime.

Published by Plain Intelligence — practical AI SEO, GEO, and technical SEO, documented in plain English. About Plain Intelligence →

↑ Back to Schema Markup · Explore all articles · Free tools & resources · Glossary