Introduction#
This tool helps you crawl documentation websites incrementally, extract their content, and create a search index in Upstash Search.
Usage#
It is available both as a CLI tool and a library.
CLI Usage#
You can run the CLI directly using npx (no installation required):
Or with command-line options:
You will be prompted for any missing options:
- Your Upstash Search URL
- Your Upstash Search token
- (Optional) Custom index name
- The documentation URL to crawl
What the Tool Does#
- Discover all internal documentation links
- Crawl each page and extract content
- Track new or obsolete data
- Upsert the new records into your Upstash Search index
Library Usage#
You can also use this as a library in your own code:
Obtaining Upstash Credentials#
- Go to your Upstash Console.
- Select your Search index. (See How to Create Search Index)
- Under the Details section, copy your
UPSTASH_SEARCH_REST_URLandUPSTASH_SEARCH_REST_TOKEN.--upstash-urlcorresponds toUPSTASH_SEARCH_REST_URL--upstash-tokencorresponds toUPSTASH_SEARCH_REST_TOKEN
Further Reading#
Try combining this tool with Qstash Schedule to keep your database up to date with docs. You may deploy your crawler on a server and call it on a schedule regularly to fetch updates in your docs. Check out our example project for implementation details: A modern documentation library to search and track the docs.
For further insights, see @upstash/search-crawler