Archiving Web Data Big and Small: Wayback, Webrecorder, CommonCrawl
Details
Topic:
This talk will give a broad overview of the field of web archiving and its various forms, from large scale web crawling to high fidelity web recording.
We will cover how web data is crawled, stored and accessed by such existing services such as the Wayback Machine, CommonCrawl and the new Webrecorder project ( https://webrecorder.io/ ), and what 'web archiving' even means. The talk will cover the ISO standard WARC format, web archive index APIs for CommonCrawl and Wayback Machine, and high-fidelity web archiving with Webrecorder, and present an overview of various open source tools and technologies available for working with web archives. Current and future challenges facing web archiving technology will also be discussed.
Speaker:
Ilya Kreymer has been worked in the field of web archiving for several years. He currently leads the development of the https://webrecorder.io (https://webrecorder.io/) in collaboration with Rhizome, a digital arts non-profit based in New York.
He has contributed to various open source web archiving and open access projects, including CommonCrawl project, Hypothesis annotation project and is also developing a new web archive replay and access toolset ( https://github.com/ikreymer/pywb ). Previously, he has worked at the Internet Archive focusing on web archive replay and access, and led the development of several new features and APIs for the Internet Archive Wayback Machine. Ilya is also a recipient of the Shuttleworth Foundation Flash Grant ( https://www.shuttleworthfoundation.org/flashgrants/ ) for Nov 2015.
Last week, Ilya released http://oldweb.today , a old browser emulator that allows users to browser old websites from various archives
