crawling github #209232
🏷️ Discussion TypeQuestion 💬 Feature/Topic AreaAPI BodyHi my name is Joshua an i don't know if this fits into this category but anyway, i'v searched google how to get permission to crawl (index) github to my own search engine (SPS search) and i cant find a answer |
Replies: 3 comments 4 replies
AnswerGitHub does not require a special "crawling permission" to index public repositories. For an indexer/search engine, the recommended approach is to use GitHub's REST or GraphQL API instead of scraping GitHub pages. API limits
What permissions do I need?It depends on what you're indexing:
So, there is no separate permission that simply enables "GitHub crawling." Use the GitHub API and request only the permissions your application actually needs. If this answer was helpful, please mark it as answered. |
|
The For what you're trying to build, though, there are probably better ways than crawling the whole site:
curl -H "Authorization: Bearer YOUR_TOKEN" \
https://api.github.com/repos/OWNER/REPO/readmeThe README endpoint also works without authentication for public repositories.
So for SPS Search, I'd use the API and public source repositories wherever possible. For pages that you specifically need to crawl from |
|
If your goal is to build a search engine that indexes GitHub content, I wouldn’t try to bypass the "robots.txt" restrictions. For public repositories, you can use GitHub’s REST or GraphQL API to discover repositories and retrieve things like READMEs and repository files. For GitHub Docs specifically, you can also work from the public "github/docs" repository instead of crawling the documentation website directly. For example, you can search for repositories through the GitHub API, then use the repository contents endpoints to retrieve the files you want to index. Caching the results will also help you avoid unnecessary API requests and rate-limit problems. If SPS Search needs pages that GitHub specifically blocks from crawling, the safest route is to contact GitHub Support rather than trying to work around "robots.txt". |
The
blocked by robots.txtmessage basically means your crawler is respecting GitHub's rules. You shouldn't try to bypass those blocks. GitHub's currentrobots.txttells crawlers that want to crawl GitHub to contact GitHub Support. ([github.com](https://github.com/robots.txt))For what you're trying to build, though, there are probably better ways than crawling the whole site:
GitHub Docs: The documentation is available in the
github/docsrepository, so you can clone that repo and index the Markdown files instead of crawlingdocs.github.com. Just check the repository's license before redistributing the content.Repository READMEs/files: Use the GitHub REST API. For example:
curl -H "A…