All pages
The ways in

From a URL

Point the newsroom at a page and it writes a new article out of it - the page is a priority source, never a template, and none of its prose is copied.

Point the desk at a URL and it writes a new article out of it. The page is handed over as a priority source, not as a template: the nine desks research around it, and nothing of its prose is copied.

One page is read first, a list is not

With a single URL you see what we actually got before spending a generation. With several we skip that step - twenty fetches up front is a minute of staring at spinners, and the run fetches each page anyway, so a page we cannot read fails only its own row.

One URL, or a list

Press Enter for a second line and it is a run, the same gesture that turns one topic into an assignment. Delete back to one line and the read step returns.

https://example.com/one
https://example.com/two
https://example.com/three

Paste or upload a list you already have

A whole column out of a spreadsheet, or upload .txt .csv .xlsx .docx .pdf .json, several files at once. Any line is scanned for its http(s) address and the rest of the line is treated as a label, so a Title, URL export works as it stands. Excel is read from its first column. A line with no address in it is named rather than quietly dropped.

Upload also reads a sitemap. Give it yoursite.com/sitemap.xml and it lists what the site publishes, optionally narrowed to a path like /blog/ and to pages not touched in the last 3, 6 or 12 months, taking the first N. It reports three numbers - how many the sitemap lists, how many survived the filters, how many are being taken - because a filter that quietly ate the list looks exactly like one that worked.

Nothing is sent unread

One URL or twenty-five, the pages are fetched and shown to you first: the headline we found, the domain, the word count, the date. A page we could not read is named and left out of the run rather than generated against.

For a list this is a checklist. Every page arrives ticked, one selection decides both what gets read and what gets sent, and it survives the read - so pruning forty-eight rows is not work you do twice.

This catches three things a URL alone cannot: a dead link, a paywall, and a section index or author page, which a sitemap import will happily hand over and which is a list of headlines rather than a piece of writing.

A page we cannot read says which

Paywalled, behind a login, or drawn entirely in the browser all extract nothing, and each asks something different of you. The screen names the one it hit instead of writing an article about a page it never read.

Every page is a full article

Ten URLs is ten generations, not one cheap batch, and a run is capped at 25. Repeats are dropped so the same page is never paid for twice.

Related: Dossier when you want it written from your links and nothing else · Assignment