Skip to content
Agent Search Engine.

Open source · Python

Best Python research & data agents

The 19 open-source research & data agents written in Python in this index, ranked by real maintained adoption (GitHub stars and recent commit activity), never by sponsorship. Scrapling leads the set at 84,252 GitHub stars, and 15 of the 19 have shipped a commit in the last 90 days. Every project here is Python-first, self-hostable, and free to run; the trade-off is you host and maintain it yourself. Right now it's led by Scrapling (84k stars), with TrendRadar and BettaFish close behind. Rankings shift as projects gain stars and ship commits.

19 open-source records · all research & data agents · more best-of lists · how we rank

ScraplingInfrastructureAn adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!Open sourcePython · BSD-3-ClauseVerified · Active 84kGitHub starsTrendRadarAgentAI-driven public opinion & trend monitor with multi-platform aggregation, RSS, and smart alerts.Open sourcePython · GPL-3.0Verified · Active 63kGitHub starsBettaFishAgentA multi-agent public-opinion analysis assistant anyone can use -- reconstructs the real shape of public sentiment and forecasts where it is heading. Built from scratch, no framework dependencies.Open sourcePython · GPL-2.0Verified · Active 42kGitHub starskhojAgentYour AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research.Open sourcePython · AGPL-3.0Verified 38kGitHub starsstormAgentAn LLM-powered knowledge curation system that researches a topic and generates a full-length report with citations.Open sourcePython · MITVerified 32kGitHub starsgpt-researcherAgentAn autonomous agent that conducts deep research on any data using any LLM providersOpen sourcePython · Apache-2.0Verified · Active 30kGitHub starshaystackFrameworkOpen-source AI orchestration framework for building context-engineered, production-ready LLM applications.Open sourcePython · Apache-2.0Verified · Active 27kGitHub starsDeepResearchAgentTongyi Deep Research, the Leading Open-source Deep Research AgentOpen sourcePython · Apache-2.0Verified 20kGitHub starslocal-deep-researchAgent~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google,...). 10+ search engines - arXiv, PubMed, your private documents.Open sourcePython · MITVerified · Active 9.1kGitHub starsMiroThinkerAgentMiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.Open sourcePython · Apache-2.0Verified 8.4kGitHub starsdeep-searcherAgentOpen Source Deep Research Alternative to Reason and Search on Private Data. Written in Python.Open sourcePython · Apache-2.0Verified · Active 8.3kGitHub starsunstractPlatformLLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline WorkflowsOpen sourcePython · AGPL-3.0Verified · Active 7.3kGitHub stars
For agent buildersMake your product part of the discovery.Explore advertising

Frequently asked

What are the best Python research & data agents?
By maintained adoption — GitHub stars plus recent commit activity — Scrapling, TrendRadar and BettaFish lead the Python research & data agents in this index of 19. The full ranking is below; it reflects what is genuinely used and still maintained, not what pays.
Are these Python research & data agents open source and free?
Yes — every project on this page is open source and written primarily in Python. Apache-2.0 is the most common licence here, on 9 of 19, across 6 licences in total. The software is free to run and self-host; you still pay for the infrastructure you run it on and any model or API usage it makes.
Why choose Python research & data agents specifically?
Staying in your team's primary language — Python — makes self-hosting, extending, and debugging far easier, because you can read and modify the source directly instead of treating it as a black box. It also makes upstream activity something you can judge: 15 of these 19 shipped in the last 90 days, and reading a quiet project's source is how you tell finished from abandoned.