
DSpace
TinyCat
Greenstone Digital Library
Invenio
FOLIO
FedoraCommons
BiblioteQ
Evergreen ILS
Diggernaut
import.io
Octoparse
artoo.js
Webhose.io
Crawlbase
eScraper
Agenty
Company offering cloud based web scraping and data extraction platform that works not only with HTML pages as data source but also with JS, JSON, XML, documents like iCal, XSLX, XLS, CSV and images. Extracted data kept in the database as dataset which can be downloaded in various formats, retrieved via API or pushed to any other destination upon completion. Integrated with such services like Zapier, Tableau, OSM, Luminati, DeathByCaptcha.
DSpace
DiggernautTinyCat - The online catalog and integrated library system for tiny libraries, powered by LibraryThing.
import.io - Import. io helps its users find the internet data they need, organize and store it, and transform it into a format that provides them with the context they need.
Greenstone Digital Library - Greenstone is a suite of software tools for building and distributing digital library collections...
Octoparse - Octoparse provides easy web scraping for anyone. Our advanced web crawler, allows users to turn web pages into structured spreadsheets within clicks.
Invenio - Invenio is a free, open-source software to run a digital library or document repository on the web.
artoo.js - Artoo.js provides script that can be run from your browserโs bookmark bar to scrape a website and return the data in JSON format.