For the complete documentation index, see llms.txt. This page is also available as Markdown.

4CAT

4CAT is a tool designed for the easy collection and analysis of online datasets. It allows researchers to uncover patterns and trends in data from social media and other digital platforms.

URL

https://4cat.nl/ (Latest release: v1.55, 2026-06-22; last checked: 2026-06-26)

Description

4CAT is an open-source, containerized web application (commonly deployed via Docker) for capturing, importing, and analyzing online data through an accessible browser interface. Researchers create datasets from supported platforms or import datasets collected with companion tools; they can then run a large library of “processors” to explore trends, content, networks, and media. 4CAT is designed to make repeatable capture and analysis workflows available without requiring programming for day-to-day use, but installation and maintenance do require technical setup.

Features

Data Sources

Direct capture in 4CAT (actively supported in the project README)

  • 4chan and 8kun

  • Bluesky

  • Telegram

  • TikTok (from a list of TikTok post URLs)

  • Tumblr

Import via Zeeschuimer (browser-based capture, then import to 4CAT)

  • TikTok (posts and comments)

  • Instagram (posts only)

  • X/Twitter

  • LinkedIn

  • 9gag

  • Imgur

  • Douyin

  • Gab

  • Truth Social

  • Threads

  • Pinterest

  • RedNote/Xiaohongshu

Import from other tools or files

  • Facebook and Instagram: via Facepager exports (CrowdTangle is discontinued; only legacy exports apply if already obtained)

  • YouTube videos and comments: via YouTube Data Tools

  • Weibo: via Bazhuayu

  • Generic imports: CSV files and common media formats (audio, video, images) can also be uploaded for analysis.

Scheduling

Once data sources are configured, 4CAT can be used in ongoing workflows (repeated collection + scheduled processor runs), depending on the connector and your deployment setup.

Processors

Processors are 4CAT’s built-in tools for working with a dataset after it has been collected or imported. They can clean, filter, transform, visualise, export, or analyse the data. For example, a processor might create a frequency chart, show activity over time, extract terms, prepare a network export, or run a more advanced text-analysis workflow. In short: data sources get data into 4CAT; processors help you turn that data into something you can inspect, interpret, or export.

  • Filtering and transformation: filter by value/date/keywords; anonymise fields; convert between common formats (CSV, JSON, NDJSON)

  • Metrics and exploration: counts and distributions over time; “top terms” style summaries; thread and post metrics

  • Text analysis: entity extraction, topic modeling, word counts, and other NLP-style processors

  • Networks: exports and processors for network analysis and visualization (including GEXF outputs)

  • Media analysis: image walls, media downloads, and image-oriented processors; newer releases also add processors that support LLM-assisted annotation and evaluation, depending on configuration.

Examples

The example below shows creating a new dataset and then visualizing results with a stream graph (a stacked time-series view that helps you compare how topics/terms rise and fall over time).

In this example, Tumblr is selected to collect posts/comments. Here, the dataset is defined via tags entered in the Tags/blogs field (e.g., #liminalspaces), and the results are visualized over time to compare term/topic trends.

[Screenshot: “Create dataset” screen showing Tumblr selected + Tags/blogs field filled]
4CAT’s “Create new dataset” screen for Tumblr. This example searches by tag, using liminal spaces, so 4CAT retrieves posts explicitly tagged with that phrase rather than generic keyword matches. The query is intentionally narrow to reduce load on the shared instance and avoid large, slow collections. Before creating the dataset, author information is pseudonymised/replaced and “Make dataset private” is enabled. Use the date range fields to narrow the collection further, then check the output for missing time periods or gaps.

Example of the customization of a steam graph visualization in 4Cat.

Example of a 4CAT stream graph. Each colour is a selected term/category, and thicker bands mean more matching posts/items in that time period. The large peak near late summer 2021 marks a short burst of activity worth inspecting in the underlying dataset. Downward bands are not negative values; stream graphs are centred, so the important signal is band thickness over time.

Cost

Hosting cost may apply if 4Cat is deployed as a web application for a team.

Level of difficulty

starstarstarstar

Installation/server maintenance (often via Docker) is the main hurdle; day-to-day use in the web UI is typically easier once an instance is running.)

Requirements

  • Platforms: runs on Linux, Windows and macOS.

  • Docker: the application has been containerised with Docker so Docker needs to be installed for the application to run.

  • Memory Requirements: 16GB of RAM

  • API key: Some of the custom datasets may require an API key to be configured.

For more information on hardware requirements see: https://github.com/digitalmethodsinitiative/4cat/wiki/What-hardware-do-I-need-to-run-4CAT%3F

Limitations

  • Technical Installation: the primary limitation of 4Cat is that to install it you need specialised technical knowledge of tools like Docker and a server to run it on.

  • Data Access: coverage and depth vary by platform. The project notes that some built-in sources are untested or require special API access, and platform policy/technical changes can reduce what is collectible.

  • Complex Queries: Users with limited technical expertise may find it challenging to construct complex queries or fully utilize the tool's capabilities without a steep learning curve.

  • Processing Time: Large datasets or complex analysis tasks may require significant processing time, which could impact efficiency for time-sensitive research.

  • Update Frequency: Platform coverage can change quickly. The 4CAT project notes that some built-in platform support is untested or requires special API access, and platform policy/technical changes can break capture. Treat the platform list as “as of last checked” and verify against the README/wiki before relying on it for time-sensitive work.

  • Cost: While 4CAT itself may be free to use, certain analyses may require substantial computational resources or access to premium data sources, which may incur costs.

  • Rate Limits: some services throttle collection. In 4CAT’s own docs, some sources are explicitly documented as hard to scrape “within 4CAT itself” due to aggressive rate limiting (e.g., TikTok comments, Imgur), with a recommendation to import data collected elsewhere.

Ethical Considerations

When using 4CAT for research, several ethical considerations must be taken into account:

  • Privacy and Consent: Researchers must navigate the complex landscape of user privacy, especially when collecting data from social media platforms. It is crucial to ensure that data collection complies with platform privacy policies and respects users' consent, especially when users have not explicitly agreed to share their data for research purposes.

  • Data Anonymization: Ensuring data is anonymized to protect individuals' identities is paramount. This involves removing or obfuscating any identifiable information before analysis or publication of the research findings.

  • Bias and Representation: The tool's reliance on accessible platform data may introduce bias, as not all voices and perspectives are equally represented online. Researchers should be aware of these limitations and consider them when drawing conclusions from their data.

  • Impact on Subjects: There should be careful consideration of the potential impact of the research on the subjects being studied, especially if the findings could adversely affect them or their communities.

  • Compliance with Legal Standards: Ensuring adherence to applicable laws and regulations, such as the GDPR in the European Union, is essential. Researchers must be mindful of the legal implications of data collection, storage, and analysis practices.

Guides and articles

To effectively use 4Cat, especially for beginners or those looking to refine their skills, the following resources are highly recommended:

Official Wiki

Tutorials and Articles

  • 4CAT exercises (no date). Available via the official 4cat.nl “Exercises” link (Accessed: 2026-01-31).

  • ‘CAT4SMR – Capture and Analysis Tools for Social Media Research’ (no date). Available at: https://cat4smr.humanities.uva.nl/ (Accessed: 13 May 2024).Peeters, S. and Hagen, S. (2022) ‘The 4CAT Capture and Analysis Toolkit: A Modular Tool for Transparent and Traceable Social Media Research’, Computational Communication Research, 4(2), pp. 571–589.

  • "As a researcher, this tool saves me a lot of work and stress” - News - Utrecht University (2023). Available at: https://www.uu.nl/en/news/as-a-researcher-this-tool-saves-me-a-lot-of-work-and-stress (Accessed: 13 May 2024).

Video Tutorials

Developer Resources

Community and Support

Tool provider

Digital Methods Initiative https://digitalmethods.net/, Netherlands

Similar Tools

Zeeschuimer - Best if you mainly need in-browser capture from hard-to-scrape platforms and do not yet need 4CAT’s processor library.

Facepager - Better for API-based or custom web/API collection workflows; less of an end-to-end analysis environment than 4CAT.

YouTube Data Tools - Better if your project is only about YouTube and you want dedicated extractors rather than a multi-platform toolkit.

Gephi - Better for advanced network exploration and visual styling after exporting GEXF or edge lists from 4CAT.

Voyant Tools - Better for lightweight browser-based text exploration when you already have a clean corpus and do not need 4CAT’s capture or import workflow.

Advertising Trackers

Page maintainer

Martin Sona

Last updated

Was this helpful?