
ScrapingBee
Apify
Scraper API
Zyte
Scrapy
Bright Data
Firecrawl
Web Scraper
BFO Java PDF Library
PDFCrowd
Document Cyborg
DocRaptor
PDFSwitch
PDFDancer
HTML PDF API
pdflayer
Web Scraping is hard, scraping at scale can be very challenging.
You have to handle:
ScrapingBee is a simple API that does all the above for you, and much more.
ScrapingBee
BFO Java PDF LibraryBased on our record, ScrapingBee seems to be more popular. It has been mentiond 3 times since March 2021. We are tracking product recommendations and mentions on various public social media platforms and blogs. They can help you identify which product is more popular and what people think of it.
If youโre worried about the security risks, edge cases, maintenance pain and scaling challenges of self hosting there are various solid hosted alternatives: - https://browserless.io - low level browser control - https://scrapingbee.com - scraping specialists - https://urlbox.com - screenshot specialists* Theyโre all profitable and have been around for years so you can depend on the businesses and the tech. *... - Source: Hacker News / over 1 year ago
If you really just need the data you can use something like https://scrapingbee.com to scrape the info from the various price pages to make sure your info is always up to date. Source: over 3 years ago
Well done! And posting here was a great idea. Not sure I would have found scrapingbee.com otherwise. We will probably become a customer. Signed up for the trial account. Source: about 4 years ago
Apify - Apify is a web scraping and automation platform that can turn any website into an API.
PDFCrowd - Pdfcrowd is a Web/HTML to PDF online service. Convert HTML to PDF online in the browser or in your PHP, Python, Ruby, .NET, Java apps via the REST API.
Scraper API - Scale Data Collection with a Simple API.
Document Cyborg - Easily save a web page as a document (PDF, WORD, EPUB, ODT, RTF or Text)
Zyte - We're Zyte (formerly Scrapinghub), the central point of entry for all your web data needs.
DocRaptor - As the only API powered by the Prince HTML-to-PDF engine, DocRaptor provides the best support for complex PDFs with powerful support for headers, page breaks, page numbers, flexbox, watermarks, accessible PDFs, and much more