Skip to content
Discussion options

You must be logged in to vote

The blocked by robots.txt message basically means your crawler is respecting GitHub's rules. You shouldn't try to bypass those blocks. GitHub's current robots.txt tells crawlers that want to crawl GitHub to contact GitHub Support. ([github.com](https://github.com/robots.txt))

For what you're trying to build, though, there are probably better ways than crawling the whole site:

  • GitHub Docs: The documentation is available in the github/docs repository, so you can clone that repo and index the Markdown files instead of crawling docs.github.com. Just check the repository's license before redistributing the content.

  • Repository READMEs/files: Use the GitHub REST API. For example:

curl -H "A…

Replies: 3 comments 4 replies

Comment options

You must be logged in to vote
4 replies
@8493834
Comment options

@8493834
Comment options

@8493834
Comment options

@8493834
Comment options

Comment options

You must be logged in to vote
0 replies
Answer selected by 8493834
Comment options

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
Apps API and Webhooks Discussions related to GitHub's APIs or Webhooks Question Ask and answer questions about GitHub features and usage Welcome 🎉 Used to greet and highlight first-time discussion participants. Welcome to the community! source:ui Discussions created via Community GitHub templates API Discussions around GitHub API platform and docs
4 participants