Digital Methods Initiative Tool Archive

This is an archive of tools developed by the DMI & affiliates in the past years. Note that not all tools are still available and while some are actively maintained, others are not.


 
4CAT App and Web Studies
4CAT extension to collect screenshots and HTML from the general web, various App stores (Apple, Google), Cloud service stores (AWS, Google Cloud, Azure), and the Way Back Machine
4CAT: Capture and Analysis Toolkit
Create datasets from a variety of web platforms - including Reddit, Telegram, 4chan and others - and analyze them.
App Tracker explorer
DMI App Tracker Tracker is a tool to detect in a set of APK files predefined fingerprints of known tracking technologies or other software libraries.
Brightbeam
Interactively capture and inspect third-party trackers encountered while browsing.
Bubble Lines
Input tags and values to produce relatively sized bubbles. Output is an svg.
Colors For Data Scientists
Generate and refine palettes of optimally distinct colors. (by Médialab Sciences-Po)
Compare Lists
Compare two lists of URLs for their commonalities and differences.
Deduplicate
Replicates the tags in a tag cloud by their value
Expand Tiny Urls
Expands URLs that have been shortened by tools like tinyurl.com or bit.ly. Often used in social media such as Twitter or Facebook.
Extract URLs
Extracts URLs from an Issuecrawler result file (.xml). Useful for retrieving starting points as well as a clean list of the actors in the network.
Geo IP
Translates URLs or IP addresses into geographical locations
Harvester
Extract URLs from text, source code or search engine results. Produces a clean list of URLs.
Image Scraper
Scrape images from a single page.
Internet Archive Wayback Machine Link Ripper
Scrapes links from the Wayback Machine
Internet Archive Wayback Machine Network Per Year
Enter a set of URLs and the archived versions closest to 1 July for a specific year are retrieved. Thereafter links are extracted and a network file is output.
Issuecrawler
Enter URLs and the Issue Crawler performs co-link analysis in one, two or three iterations, and outputs a cluster graph. The Issue Crawler also has modules for snowball crawling (up to 3 deg...
Language Detection
Detects language for given URLs. The first 1000 characters on the Web page(s) are extracted, and the language of each page is detected.
Link Ripper
Capture all internal links and/or outlinks from a page.
Raw Text to Tag Cloud Engine
Takes raw text, counts the words and returns an ordered, unordered or alphabetically ordered tagcloud.
Rip Sentences
Rip text from a specified page and force line breaks between sentences.
Robots.txt Discovery
Display a site's robot exclusion policy.
Schuifmaat
Schuifmaat is a browser extension that enables actual infinite scroll on X (fka Twitter)'s search result pages.
Source Code Search
loads a URL and searches for patterns in the page's source code
TLD counts
Enter URLS, and count the top level domains.
Table to Net
Extract a network from a table. Set a column for nodes and a column for edges. It deals with multiple items per cell. (by Médialab Sciences-Po)
Tag Cloud Combinator
Enter two or more tag clouds and the values of each tag will be summed.
Tag Cloud Generator
Input tags and values to produce a tag cloud. Output is in SVG.
Tag Cloud HTML Generator
Input tags and values in wordle format to produce a HTML tag cloud or tag list.
Tag Cloud To Wordle
This tool allows one to transform a normal tag cloud into a fancy Wordle one.
Text Ripper
Rip all non-html (i.e. text) from a specified page.
Timestamp Ripper
Rips and displays a web page's last modification date (using the page's HTML header). Beware of dynamically generated pages, where the date stamps will be the time of retrieval.
Triangulation
Enter two or more lists of URLs or other items to discover commonalities among them. Possible visualizations include a Venn Diagram.
Wikipedia Cross-Lingual Image Analysis
Makes the images of all language versions of a Wikipedia article comparable.
Wikipedia Edits Scraper and IP Localizer
Scrapes Wikipedia history and does IP to Geo for anonymous edits
Wikipedia Entry Check
This tool checks if the issues exist as a Wikipedia page, i.e., an article. If it exists it checks whether the organization is mentioned on that page.
Wikipedia History Flow Companion
This script allows you to specify a range of Wikipedia revisions for use with the History Flow visualization.
Wikipedia TOC Scraper
Scrape Table of Contents for revisions of a wikipedia page and explore the results by moving a slider to browse across chronologically ordered TOCs.
Wikipedia categories scraper
Scrape Wikipedia for the categories of articles and the categories of related articles in different languages.
YouTube Data Tools
A collection of simple tools for extracting data from the YouTube platform via the YouTube API v3.
Zeeschuimer
Zeeschuimer is a browser extension that monitors internet traffic while you are browsing a social media site, and collects data about the items you see in a platform's web interface for late...
Zoekplaatje
Zoekplaatje is a Firefox extension to facilitate the capture of search engine results.
 

Inactive and deprecated tools

The tools listed in this section are no longer actively maintained and most will not be available or are no longer functional. We keep them listed here as a reflection of the analytical focus of DMI over the years and as an illustration to the ever-changing nature of internet research.


 

Deprecated!

Actor Profiler
The Actor Profiler works in tandem with the Issue Crawler. It calculates the top ten actors on an Issue Crawler map, and profiles each of them. The profile consists of each actor's inlinks a...

Deprecated!

Amazon Book Explorer
Provides different analytics for Amazon.com's book search

Deprecated!

Amazon Related Product Graph
This command-line Python script allows you to enter a (set of) ASIN(s) and crawl its recommendations up til a user-specified depth.

Deprecated!

Censorship Explorer
Check whether a URL is censored in a particular country by using proxies located around the world.

Deprecated!

Compare Networks Over Time
Compares Issue Crawler networks over time, and displays ranked actor lists. The over time module is best used in tandem with the Issue Crawler scheduler. The results may be plotted to line g...

Deprecated!

Convert Issuecrawler to Navicrawler
Convert an Issuecrawler XML file into the WXSF format of the Navicrawler file. For Navicrawler, see http://webatlas.fr/wp/navicrawler/.

Deprecated!

Del.icio.us Related Tag Scraper
Enter a tag and show which other tags are most related to it. Visualized in a tag cloud.

Deprecated!

Del.icio.us Tags Scraper
Enter a URL and show which tags are most related to it. Visualized in a tag cloud.

Deprecated!

Discus Comment Scraper
This tool scrapes threads and comments from websites implementing the Disqus commenting system.

Deprecated!

Dorling Map Generator
Input tags and values to produce a Dorling Map (i.e. bubbles). Output is an svg.

Deprecated!

FlickrPhotoPoolNetwork
Outputs the contact network of Flickr users from a Photopool in .gdf format

Deprecated!

Geo Extraction
Extracts geographic locations from text.

Deprecated!

Github organizations meta-data lookup
Extract the meta-data of organizations on Github

Deprecated!

Github repositories meta-data lookup
Extract the meta-data of Github repositories

Deprecated!

Github repositories scraper
Scrape Github for forks of projects

Deprecated!

Github scraper
Scrape Github for user interactions and user to repository relations

Deprecated!

Github user meta-data lookup
Extract meta-data about users on Github

Deprecated!

GithubContributorsScraper
Find out which users contributed source code to Github repositories

Deprecated!

Google Autocomplete
Retrieves autocomplete suggestions from Google

Deprecated!

Google Blog Search Scraper
Batch queries Google Blog Search. Query the resonance of a particular term, or a series of terms, in a set of blogs.

Deprecated!

Google Image Scraper
Query images.google.com with one or more keywords, and/or use images.google.com to query specific sites for images.

Deprecated!

Google News Scraper

Deprecated!

Google Play Similar Apps
DMI Google Play Similar Apps is a simple tool to extract the details of individual apps, collect ‘Similar’ apps, and extract their details.

Deprecated!

Google Play Store Scraper
Google Play Store Scraper is a simple tool to extract the details of individual apps, collect their related apps, retrieve app permissions, and retrieve a list of apps for a given keyword.

Deprecated!

Google Reverse Image scraper
Scrape Google for occurance of images

Deprecated!

Google Scraper Teaser Text Ripper
Analyzes a results file from the Google Scraper to get unique phrases from the teaser (or lead) text for each google search return.

Deprecated!

Googlescraper (Search Engine Scraper)
Batch queries Google. Query the resonance of a particular term, or a series of terms, in a set of Websites.

Deprecated!

Instagram Hashtag Explorer
Retrieve either the latest media tagged with a specified term or the media around a particular location.

Deprecated!

Instagram Network
Gets the follow or follower network from a set of Instagram users.

Deprecated!

Instagram Scraper
Retrieves Instagram images for hashtags, locations, or user names.

Deprecated!

Issue Discovery Tool
Enter URLs, and discover the most relevant words and phrases contained in them. One also may enter text, or an Issuecrawler result file (.xml). Based on Yahoo Term Extraction

Deprecated!

Issue Dramaturg
Enter up to 3 URLs as well as a key word. The Issuedramaturg queries Google for the key word, and shows the Pageranks of the URLs over time. The output is a graph of the Pagerank of the URLs...

Deprecated!

Issue Feed
Enter URLs. Issuefeed.net summarizes the content of a set of sites as a list of key words.

Deprecated!

Issue Geographer
Geo-locates the organizations on an Issue Crawler map, using whois information, and visualizes the organizations' registered locations on a geographical map.

Deprecated!

Issue Network Cloud
Extracts the URLs from an Issue Crawler result, and allows the user to query all the URLs or hosts for key words (through Google). It outputs a tag cloud. Good for the analysis of the conten...

Deprecated!

Itunes Store
Queries the itunes store

Deprecated!

Keyword Resonance Tool
Queries Google for a set of keywords to calculate overlap.

Deprecated!

Like Scraper
Retrieves the number of likes on Facebook for a series of URLs

Deprecated!

Lippmannian Device
The Lippmannian device is named Walter Lippmann, and provides a coarse means of showing actor partisanship.

Deprecated!

Lippmannian Device To Gephi
This tool allows one to visualize the output of the Lippmannian device as a network with Gephi.

Deprecated!

Netvizz
Extracts various datasets from Facebook.

Deprecated!

NetvizzToSentiStrength
Adds sentiment analysis to short texts via Sentistrength

Deprecated!

News Agencies Scraper
Scrape various news agencies for particular keywords and extract titles, images, dates and full text.

Deprecated!

Open Calais
Discovers the most relevant words and phrases among a set of websites, within a text, or within an issue network. Based on Reuters Open Calais.

Deprecated!

Pagerank
Discover a website's rank in the returned Google results per issue/query.

Deprecated!

Pinterest Scraper
Scrape Pinterest for pins

Deprecated!

RSS Discovery
Discovers RSS/ATOM/RDF feeds in websites.

Deprecated!

Ranked Deep Pages from Core Issue Crawler Network
Enter an Issuecrawler XML file and this script will get out all pages from the core network and rank those by pages by inlink count.

Deprecated!

Raw Text to Tree Map Engine
Input a text or Google Scraper results file and visualize the word counts in a Tree Map. Output is in svg.

Deprecated!

Screenshot generator
Produce screenshots for a list of URLs

Deprecated!

Search Engine Scraper

Deprecated!

Surfer Issue Pathways
Building upon Alexa's related sites feature, this tool determines which sites are likely to be in the actual surfer paths of other sites related to the same issue.

Deprecated!

Technorati Actor Resonance Chart Builder
Queries Technorati and charts the resonance among a single actor and any number of related actors.

Deprecated!

ToolAlchemy
Implements a DMI interface around the AlchemyApi

Deprecated!

Tracker Tracker
DMI App Tracker Tracker is a tool to detect in a set of URLs predefined fingerprints of known web tracking technologies.

Deprecated!

Tree Map Generator
Input tags and values to produce a Tree Map. Output is in svg.

Deprecated!

Tumblr
a simple co-hashtag and post data tool for Tumblr

Deprecated!

Twitter Capture and Analysis Toolset (DMI-TCAT)
Captures tweets and allows for multiple analyses (hashtags, mentions, users, search, ...)

Deprecated!

Twitter Scraper
Very simple Twitter scraper

Deprecated!

Wikipedia Network Analysis
Find a hyperlink network around a Wikipedia topic (see here for more information, and [[http://thepoliticsofsystems.net/2011/02/21/visualizing-wikipedia-on-data-vi...

Deprecated!

Yahoo Inlink Scraper
Retrieves all inlinks to a site, according to Yahoo Site Explorer.

Deprecated!

YouTube Response Networks
Enter the URL for a YouTube movie and retrieve all video responses.

Deprecated!

YouTube Video Discovery
Analyzes a results file from the Google Scraper to discover, count, and rank YouTube and Google Video links in the descriptions. (It may also be used for Ikbis and other video compilati...

Deprecated!

iTunes App Store Scraper
The iTunes App Store Scraper is a simple tool to extract the details of individual apps, collect their related apps, and retrieve a list of apps for a given keyword.
 
Topic revision: r18 - 25 Aug 2026, StijnPeeters
This site is powered by FoswikiCopyright © by the contributing authors. All material on this collaboration platform is the property of the contributing authors.
Ideas, requests, problems regarding Foswiki? Send feedback