Based on our record, Archive-It should be more popular than Urlbox.io. It has been mentiond 3 times since March 2021. We are tracking product recommendations and mentions on various public social media platforms and blogs. They can help you identify which product is more popular and what people think of it.
This is how I do it. I send the URLs I want scraped to Urlbox[0] it renders the pages saves HTML (and screenshot and metadata) to my S3 bucket[1]. I get a webhook[2] when it's ready for me to process. I prefer to use Ruby so Nokogiri[3] is the tool I use for scraping step. This has been particularly useful when I've want to scrape some pages live from a web app and don't want to manage running Puppeteer or... - Source: Hacker News / 2 months ago
Hi there, I run urlbox.io, which is a screenshot API that allows clicking elements, waiting for elements, injecting custom JS/CSS etc. Source: over 1 year ago
Other projects include Open Library & archive-it.org. Source: over 1 year ago
The biggest hurdle in DH of the early internet and manipulating the data at scale is creating collections from archived material that fit your needs. The internet archive has a subscription-based archival tool Archive-It but the collection starts when you (or some other person) creates it. There isn't a collection for anti-vaccine material from 1996-. With Covid, many collections have been started but these will... Source: almost 2 years ago
Archive-It is a subscription web archiving service from the Internet Archive that helps organizations to harvest, build, and preserve collections of digital content. Through our user friendly web application Archive-It partners can collect, catalog, and manage their collections of archived content with 24/7 access and full text search available for their use as well as their patrons. Source: almost 2 years ago
Screenshot Machine - Screenshot machine is an online website capturing service. Creates a screenshot or thumbnail of any online web page in couple of seconds for free.
Archive.md - archive.is allows you to create a copy of a webpage that will always be up even if the original link is down
ApiFlash - ApiFlash is a powerful serverless screenshot API built with Chromium and AWS Lambda. It can easily scale to millions of screenshots per day and has an ever growing number of satisfied big clients.
Wayback Machine - Browse through over 150 billion web pages archived from 1996 to a few months ago.
Capture by Techulus - The ultimate toolkit for converting HTML to image & PDF. Pay only for what you need.No Subscription.
Perma.cc - Perma.