The best content automation tools in one place
Content tools for the pipeline: article extraction into clean text, content inventories, metadata collection and repurposing prep - publishing stays yours. 0 bots in this category.
Content listings serve the pipeline around publishing: extracting article text and metadata from pages you may read into clean structured files, inventorying a site's content, collecting headings and schema for audits, and preparing repurposing inputs. Writing, rewriting and publishing belong to you and your CMS - nothing here posts anywhere.
Expect clean-text or row-per-page output with the source URL on every record, robots and terms obligations stated per listing, and no AI rewriting unless a listing declares an AI step with your own provider key.
Every listing ships the same four launchers - a Windows .exe and .sh scripts for macOS Apple Silicon, macOS Intel, and Linux - delivered as a single raw file with no archive or extraction step. Restore the executable bit on the shell-script platforms, and clear the macOS Gatekeeper quarantine attribute before the first run. Content inventories pair naturally with the crawl tools: run one weekly and diff to see what was published, changed or removed.
Bots in this category
No bots in this category yet
Content bots: frequently asked questions
- Can a tool extract article text to a file?
- Yes - extraction listings pull the article body, title, author and date where published into rows or clean text files, one record per URL, with the source attached.
- Do any tools publish to my CMS?
- No. Publishing needs credentials and judgement; these tools produce the input files a publishing step - yours - consumes.
- Can I inventory my own site's content?
- That is the recommended first use: crawl your own property, export URLs with headings and metadata, and use the inventory for audits and migration planning.
- Does extraction rewrite or summarise content?
- Only where a listing explicitly declares an AI step, and then with your own provider key prompted at run time. Everything else copies what the page publishes, structure intact.