Skip to content

How these pages are made

Every page here is generated from an official source PDF rather than written from memory. This page explains the process, including its limits.

The steps

  1. Collect. Official PDFs are gathered by category, exactly as the publisher issues them.
  2. Detect. Each file is checked page by page for a usable text layer, and for the script it is written in.
  3. Extract. Files with a text layer are converted to Markdown. Scanned files need OCR. Urdu files are handled separately, because the engine that is best for English is not best for Urdu.
  4. Measure. Sections and amendment footnotes are counted, and each version is compared against the one before it.
  5. Publish. One page per source document, carrying its own figures, checksum and download link.

The corpus today

  • Documents published: 141
  • Source pages covered: 25,992
  • Categories: 17
  • Documents with section data extracted: 131

Choosing an extraction engine

Engine choice is decided by measurement rather than preference. On this corpus, the Markdown converter captured about 12 percent more text from English files than a plain text dump and turned contents pages into real tables. On Urdu files the same converter returned the letters in the wrong order, while a different engine returned them correctly. So the pipeline picks per file, based on the script it detects.

Known limitations

Stating these plainly is more useful than implying the process is perfect.

  • Counts are measurements, not an official tally. Headings are detected by pattern. A section formatted unusually can be missed.
  • Schedules are not parsed. Duty rates live in schedules, which are tabular. Use the source PDF for rates.
  • A missing section is not proof of repeal. When a section appears in one version and not the next, that usually means it was omitted by a later law. It can also mean the heading changed shape. Treat it as a prompt to check the source.
  • Nastaliq Urdu cannot be reliably extracted. Those PDFs store glyphs in visual order with decomposed ligatures, so extracted text comes out scrambled. Where that happens, the page says so and publishes no text from the file.
  • Subordinate legislation is separate. SROs, rules, general orders and circulars change how a provision works in practice and are not folded into these documents.

Corrections

Because pages are generated, a defect is usually in the process rather than in one page. Fixing it corrects every affected page at once. Report anything that looks wrong through thecontact page.