A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

Extracted PDF text is page-shaped noise: running headers, page numbers, a contents page that looks like headings, and numbered lists that look like numbered sections. The question is really “how do you tell a heading from everything that resembles one”. Lead with that.

  1. Ask what the index is for and what the input looks like. Will someone open the PDF at these page numbers, or read the printed ones? Are headings numbered? Ask for two or three real pages.
  2. Remove page furniture first. Lines repeated near the top or bottom of at least half the pages, compared with digits masked, are headers and footers. Set the contents page aside but keep it (numbered lines ending in a dot leader and a page number): it is the answer key for your headings.
  3. Detect headings with more than one signal. A numbered line that doesn’t end in a period is a candidate, but it only counts if its number follows the previous heading’s: 2.2 after 2.1, 3 after 2.1.4. That rejects most list items and quantities. Say where it breaks: a false heading it accepts takes the number the real one needed, so the real heading is dropped and its section goes under the wrong title; a number the source skips drops every heading after the gap. An all-caps line is weaker evidence: accept it only where it opens a page, and never NOTE or WARNING.
  4. Compute page ranges from the next heading at the same or a higher level. A section ends on the page before that heading if the heading opens its page, otherwise on the same page. A heading orphaned at a page’s foot starts its section overleaf.
  5. Emit flat JSON with a level on each entry: easy to diff, easy to nest later.

The trap is a regex for “line starting with a number” and nothing else. The page text index drill is the timed version.

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • The document has a table of contents. How would you use it instead of skipping it?
  • A numbered list item looks exactly like a heading. What in your code tells them apart, and where does it still fail?
  • Should page numbers be the PDF’s page index or the numbers printed on the pages?
  • How would you know the index is right on a thousand real documents you can’t read?

Where answers go wrong

  • Matching any line that starts with a number as a heading, so numbered list items, years and quantities become sections.
  • Ending every section on the page where it starts, so a section that runs across three pages is indexed as one.
  • Leaving running headers and footers in, so the document title is detected as a heading on every page.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Model answer

“I’ll take a list of page strings, one per PDF page, and return sections with a title, number, level and a start and end page. Pages are the PDF’s own index, starting at one, because that’s what a viewer opens; if you want printed page numbers, I’d read them from the footer before throwing it away.