Home → Methods
Sourced, shaped, seen. Three stages, in that order, and nothing skips one. This page exists so you can judge our numbers rather than trust them — including the parts where the data is thinner than anyone would like.
We start from the primary release, never a summary of it: the ministry's own weekly bulletin, the survey report, the administrative dashboard's export.
Cleaning is limited to what makes a dataset joinable and readable. We do not improve the data.
Each topic is published in whichever form matches its shape — a map where geography is the story, a ranked ladder where comparison is, a time series where the trend is.
An official source is not the same thing as a current one. Several indicators on this platform rest on survey rounds that are years old, because no newer round exists — the most-cited household survey in India covers 2019–21, and the migration baseline is the 2011 Census.
So every figure we publish states the year it refers to, separately from the year it was published. A report released in 2023 measuring 2019–21 is labelled as measuring 2019–21. Anything else invites a reader to treat five-year-old data as today's.
If we cannot establish which period a figure refers to, we do not publish the figure. A number without a date is not evidence; it is decoration.
This is the constraint most data platforms hide. Whether a page can update itself depends entirely on how its source publishes, and Indian official data publishes in three quite different ways. We grade every source accordingly, and the grade appears on each topic page.
| Tier | What it means | How we keep it current |
|---|---|---|
| A | A machine-readable API exists and is open to us. | Polled automatically |
| B | A file — spreadsheet, CSV or structured PDF — is published at a stable address each cycle. | Parsed on a schedule |
| C | The figure exists only inside a report or behind a portal, with no stable file to fetch. | Read and logged by hand |
Most Indian official sources are tier B or C. That is the honest ceiling on how live this platform can be, and it is why each topic states its own update cadence instead of the site making one blanket promise it cannot keep.
When a source passes its expected release date without our having updated it, the topic is flagged as overdue rather than left to look current. A stale figure that admits its age is safe. A fresh-looking figure that is quietly four years old is not.
Errors get fixed, and the fix is disclosed. If you find one, write to hello@dataspeaks.global with the page and the figure in question. Corrections to a published number are noted at the foot of that page with the date they were made, so the record of what changed stays visible.
Corrections to a live figure are treated as urgent. Everything else gets a reply within three working days.
One new topic publishes each morning. Live dashboards built on weekly releases refresh with those releases, and the last-updated date always appears on the page.
Briefs for all seventeen Sustainable Development Goals are published; dashboards are being added in goal order, with requested datasets given priority. A topic moves from brief to dashboard only when its data build is complete — never partially, and never with the gaps filled in by us.
Covering all seventeen goals shallowly would be easy and useless. Each goal on DataSpeaks gets a single question worth answering properly, anchored to a named release with its year stated. Once every goal carries a published dashboard, a second topic per goal follows.
Everything above exists so that you can check our work. Every topic page names its sources and links to them, so you can go to the release yourself and see whether we read it correctly. Start with the seventeen goals.
DataSpeaks · dataspeaks.global · Methods