4CAT
4CAT is a tool designed for the easy collection and analysis of online datasets. It allows researchers to uncover patterns and trends in data from social media and other digital platforms.
URL
https://4cat.nl/ (Latest release: v1.55, 2026-06-22; last checked: 2026-06-26)
Description
4CAT is an open-source, containerized web application (commonly deployed via Docker) for capturing, importing, and analyzing online data through an accessible browser interface. Researchers create datasets from supported platforms or import datasets collected with companion tools; they can then run a large library of “processors” to explore trends, content, networks, and media. 4CAT is designed to make repeatable capture and analysis workflows available without requiring programming for day-to-day use, but installation and maintenance do require technical setup.
Features
Data Sources
Direct capture in 4CAT (actively supported in the project README)
4chan and 8kun
Bluesky
Telegram
TikTok (from a list of TikTok post URLs)
Tumblr
Import via Zeeschuimer (browser-based capture, then import to 4CAT)
TikTok (posts and comments)
Instagram (posts only)
X/Twitter
LinkedIn
9gag
Imgur
Douyin
Gab
Truth Social
Threads
Pinterest
RedNote/Xiaohongshu
Import from other tools or files
Facebook and Instagram: via Facepager exports (CrowdTangle is discontinued; only legacy exports apply if already obtained)
YouTube videos and comments: via YouTube Data Tools
Weibo: via Bazhuayu
Generic imports: CSV files and common media formats (audio, video, images) can also be uploaded for analysis.
Scheduling
Once data sources are configured, 4CAT can be used in ongoing workflows (repeated collection + scheduled processor runs), depending on the connector and your deployment setup.
Processors
Processors are 4CAT’s built-in tools for working with a dataset after it has been collected or imported. They can clean, filter, transform, visualise, export, or analyse the data. For example, a processor might create a frequency chart, show activity over time, extract terms, prepare a network export, or run a more advanced text-analysis workflow. In short: data sources get data into 4CAT; processors help you turn that data into something you can inspect, interpret, or export.
Filtering and transformation: filter by value/date/keywords; anonymise fields; convert between common formats (CSV, JSON, NDJSON)
Metrics and exploration: counts and distributions over time; “top terms” style summaries; thread and post metrics
Text analysis: entity extraction, topic modeling, word counts, and other NLP-style processors
Networks: exports and processors for network analysis and visualization (including GEXF outputs)
Media analysis: image walls, media downloads, and image-oriented processors; newer releases also add processors that support LLM-assisted annotation and evaluation, depending on configuration.
Examples
The example below shows creating a new dataset and then visualizing results with a stream graph (a stacked time-series view that helps you compare how topics/terms rise and fall over time).
In this example, Tumblr is selected to collect posts/comments. Here, the dataset is defined via tags entered in the Tags/blogs field (e.g., #liminalspaces), and the results are visualized over time to compare term/topic trends.
liminal spaces, so 4CAT retrieves posts explicitly tagged with that phrase rather than generic keyword matches. The query is intentionally narrow to reduce load on the shared instance and avoid large, slow collections. Before creating the dataset, author information is pseudonymised/replaced and “Make dataset private” is enabled. Use the date range fields to narrow the collection further, then check the output for missing time periods or gaps.Example of the customization of a steam graph visualization in 4Cat.
Cost
Hosting cost may apply if 4Cat is deployed as a web application for a team.
Level of difficulty
Installation/server maintenance (often via Docker) is the main hurdle; day-to-day use in the web UI is typically easier once an instance is running.)
Requirements
Platforms: runs on Linux, Windows and macOS.
Docker: the application has been containerised with Docker so Docker needs to be installed for the application to run.
Memory Requirements: 16GB of RAM
API key: Some of the custom datasets may require an API key to be configured.
For more information on hardware requirements see: https://github.com/digitalmethodsinitiative/4cat/wiki/What-hardware-do-I-need-to-run-4CAT%3F
Limitations
Technical Installation: the primary limitation of 4Cat is that to install it you need specialised technical knowledge of tools like Docker and a server to run it on.
Data Access: coverage and depth vary by platform. The project notes that some built-in sources are untested or require special API access, and platform policy/technical changes can reduce what is collectible.
Complex Queries: Users with limited technical expertise may find it challenging to construct complex queries or fully utilize the tool's capabilities without a steep learning curve.
Processing Time: Large datasets or complex analysis tasks may require significant processing time, which could impact efficiency for time-sensitive research.
Update Frequency: Platform coverage can change quickly. The 4CAT project notes that some built-in platform support is untested or requires special API access, and platform policy/technical changes can break capture. Treat the platform list as “as of last checked” and verify against the README/wiki before relying on it for time-sensitive work.
Cost: While 4CAT itself may be free to use, certain analyses may require substantial computational resources or access to premium data sources, which may incur costs.
Rate Limits: some services throttle collection. In 4CAT’s own docs, some sources are explicitly documented as hard to scrape “within 4CAT itself” due to aggressive rate limiting (e.g., TikTok comments, Imgur), with a recommendation to import data collected elsewhere.
Ethical Considerations
When using 4CAT for research, several ethical considerations must be taken into account:
Privacy and Consent: Researchers must navigate the complex landscape of user privacy, especially when collecting data from social media platforms. It is crucial to ensure that data collection complies with platform privacy policies and respects users' consent, especially when users have not explicitly agreed to share their data for research purposes.
Data Anonymization: Ensuring data is anonymized to protect individuals' identities is paramount. This involves removing or obfuscating any identifiable information before analysis or publication of the research findings.
Bias and Representation: The tool's reliance on accessible platform data may introduce bias, as not all voices and perspectives are equally represented online. Researchers should be aware of these limitations and consider them when drawing conclusions from their data.
Impact on Subjects: There should be careful consideration of the potential impact of the research on the subjects being studied, especially if the findings could adversely affect them or their communities.
Compliance with Legal Standards: Ensuring adherence to applicable laws and regulations, such as the GDPR in the European Union, is essential. Researchers must be mindful of the legal implications of data collection, storage, and analysis practices.
Guides and articles
To effectively use 4Cat, especially for beginners or those looking to refine their skills, the following resources are highly recommended:
Official Wiki
Tutorials and Articles
4CAT exercises (no date). Available via the official 4cat.nl “Exercises” link (Accessed: 2026-01-31).
‘CAT4SMR – Capture and Analysis Tools for Social Media Research’ (no date). Available at: https://cat4smr.humanities.uva.nl/ (Accessed: 13 May 2024).Peeters, S. and Hagen, S. (2022) ‘The 4CAT Capture and Analysis Toolkit: A Modular Tool for Transparent and Traceable Social Media Research’, Computational Communication Research, 4(2), pp. 571–589.
"As a researcher, this tool saves me a lot of work and stress” - News - Utrecht University (2023). Available at: https://www.uu.nl/en/news/as-a-researcher-this-tool-saves-me-a-lot-of-work-and-stress (Accessed: 13 May 2024).
Video Tutorials
4CAT Tutorial - Creating a Dataset - YouTube (no date). Available at: https://www.youtube.com/watch?v=VZH9SQM3dmI&list=PLWukutaRyIn31H0uPfkYlmbWvo83PnXXo&index=2 (Accessed: 13 May 2024).
4CAT Tutorial: Analyzing a dataset using processors (2021). Available at: https://www.youtube.com/watch?v=XIpGt3uzqNQ (Accessed: 13 May 2024).
4CAT Tutorial: Installing via Docker - YouTube (no date). Available at: https://www.youtube.com/watch?v=oWsB7bvNfOY (Accessed: 13 May 2024).
Developer Resources
Community and Support
Tool provider
Digital Methods Initiative https://digitalmethods.net/, Netherlands
Similar Tools
Zeeschuimer - Best if you mainly need in-browser capture from hard-to-scrape platforms and do not yet need 4CAT’s processor library.
Facepager - Better for API-based or custom web/API collection workflows; less of an end-to-end analysis environment than 4CAT.
YouTube Data Tools - Better if your project is only about YouTube and you want dedicated extractors rather than a multi-platform toolkit.
Gephi - Better for advanced network exploration and visual styling after exporting GEXF or edge lists from 4CAT.
Voyant Tools - Better for lightweight browser-based text exploration when you already have a clean corpus and do not need 4CAT’s capture or import workflow.
Advertising Trackers
Martin Sona
Last updated
Was this helpful?