How the extraction works
Your own pages, organised into reference content
- 01
Confirm the site and read robots
The address is validated as a public destination, then the site's own robots.txt rules are read and followed for every request that follows.
- 02
Discover a bounded set of pages
Links on the homepage and, where published, the sitemap produce candidate URLs. Discovery itself is capped at a couple of requests so listing pages never becomes a crawl.
- 03
Read only the pages you choose
Selected pages are fetched a couple at a time. Each URL is re-checked against the site boundary and the robots rules before it is read.
- 04
Organise what was actually there
Headings become sections, with lists, tables, existing FAQ pairs and contact links kept as written. Nothing is summarised, inferred or generated.
- 05
Review, edit and export
Every section can be rewritten, excluded or restored, and each keeps the page it came from. Exports reflect the current selection, not the original extraction.
Reference content, not training data
This produces a structured copy of text that is already published on the pages you select. It is not model training or fine-tuning data, it makes no claim that any particular assistant platform will accept it unchanged, and it never adds facts, prices, policies or answers that were not on the page. Pages built by JavaScript cannot be read, because no browser is run.
Use only public destinations you are authorized to check. By using this tool, you agree to the Terms of Use.