Home → Methods

How a dataset
becomes a chart.

Sourced, shaped, seen. Three stages, in that order, and nothing skips one. This page exists so you can judge our numbers rather than trust them — including the parts where the data is thinner than anyone would like.

The three stages

Stage 01

Sourced

We start from the primary release, never a summary of it: the ministry's own weekly bulletin, the survey report, the administrative dashboard's export.

  • For every dataset we record the publishing agency, the exact release, the reference period, the release date and the licence.
  • Secondary reporting is used only to locate a release — never as the figure itself.
  • Where two agencies publish different numbers for the same thing, we publish both and say which is which, rather than picking the convenient one.
Stage 02

Shaped

Cleaning is limited to what makes a dataset joinable and readable. We do not improve the data.

  • Place codes. States, districts and urban local bodies are mapped to standard codes so releases from different agencies can be read together. Boundary changes are tracked, because districts split.
  • Periods. Financial years, calendar years, survey rounds and rolling weeks are labelled explicitly. We never silently convert one into another.
  • Units. Reported units are preserved. Where we convert, both the original and the converted value appear.
  • Gaps. Missing values stay missing. No interpolation, no carry-forward, no filling a map so it looks finished.
  • Totals. Published components are checked against published totals. When they fail to reconcile, the discrepancy is shown on the page.
Stage 03

Seen

Each topic is published in whichever form matches its shape — a map where geography is the story, a ranked ladder where comparison is, a time series where the trend is.

  • Every page carries the source, the reference period and the release date on its face, not in a footnote.
  • Caveats that materially change how a number should be read sit next to the number.
  • Definitions travel with the figure. Forest cover, unemployment, poverty and clean water all mean something specific in the dataset that measures them.

Every figure carries its year

An official source is not the same thing as a current one. Several indicators on this platform rest on survey rounds that are years old, because no newer round exists — the most-cited household survey in India covers 2019–21, and the migration baseline is the 2011 Census.

So every figure we publish states the year it refers to, separately from the year it was published. A report released in 2023 measuring 2019–21 is labelled as measuring 2019–21. Anything else invites a reader to treat five-year-old data as today's.

The vintage rule

If we cannot establish which period a figure refers to, we do not publish the figure. A number without a date is not evidence; it is decoration.

How fresh a topic can be

This is the constraint most data platforms hide. Whether a page can update itself depends entirely on how its source publishes, and Indian official data publishes in three quite different ways. We grade every source accordingly, and the grade appears on each topic page.

TierWhat it meansHow we keep it current
A A machine-readable API exists and is open to us. Polled automatically
B A file — spreadsheet, CSV or structured PDF — is published at a stable address each cycle. Parsed on a schedule
C The figure exists only inside a report or behind a portal, with no stable file to fetch. Read and logged by hand

Most Indian official sources are tier B or C. That is the honest ceiling on how live this platform can be, and it is why each topic states its own update cadence instead of the site making one blanket promise it cannot keep.

When a source passes its expected release date without our having updated it, the topic is flagged as overdue rather than left to look current. A stale figure that admits its age is safe. A fresh-looking figure that is quietly four years old is not.

What we will not do

×Forecast, model or estimate values that an agency has not measured.
×Rank places on a composite index of our own invention.
×Interpolate across a gap so that a line looks continuous.
×Round a number in the direction that makes a point.
×Truncate an axis to exaggerate a change.
×Publish a figure whose source we cannot name.
×Use a placeholder number to fill a chart that has no data behind it yet.

Corrections

Errors get fixed, and the fix is disclosed. If you find one, write to hello@dataspeaks.global with the page and the figure in question. Corrections to a published number are noted at the foot of that page with the date they were made, so the record of what changed stays visible.

Corrections to a live figure are treated as urgent. Everything else gets a reply within three working days.

Publishing schedule

One new topic publishes each morning. Live dashboards built on weekly releases refresh with those releases, and the last-updated date always appears on the page.

Briefs for all seventeen Sustainable Development Goals are published; dashboards are being added in goal order, with requested datasets given priority. A topic moves from brief to dashboard only when its data build is complete — never partially, and never with the gaps filled in by us.

Why one topic per goal

Covering all seventeen goals shallowly would be easy and useless. Each goal on DataSpeaks gets a single question worth answering properly, anchored to a named release with its year stated. Once every goal carries a published dashboard, a second topic per goal follows.

Judge the numbers, don't trust them

Everything above exists so that you can check our work. Every topic page names its sources and links to them, so you can go to the release yourself and see whether we read it correctly. Start with the seventeen goals.

DataSpeaks · dataspeaks.global · Methods