GitHub AI Tools Intelligence

Collect AI tools and categories from GitHub Topics and Marketplace, retain source runs and sharing boundaries, then publish a browsable dashboard with a daily refresh.

3 tables126 entriesdaily refresh
See how it was built
Real workflow

From one scrape to intelligence refreshed at 09:00

Dataset creation, app publishing, email-specific sharing, and Managed Refresh connect within one agent conversation.

From one scrape to intelligence refreshed at 09:00

Help me scrape today’s recommended AI tools and categories on GitHub.

Done. I scraped GitHub live on August 4, 2026 and saved the results as local CSV/JSON snapshots, along with the scraper script.

It captured 126 tool/repo rows and 34 Marketplace categories, with no fetch errors.

Sources scraped:

A few AI tools surfaced today: LovableBot, OpenCode, CodeRabbit, cto.new, Qodo, Claude Code Action Official, AI Code Review Action, browser-use/browser-use, private-gpt, llamafile, firecrawl, langchain, and gemini-cli.


Before I create the write ticket, please confirm this placement.

Planned new GoalfyData dataset:
GitHub AI Tools Intelligence

Planned tables:

  • github_ai_tools: 126 rows from GitHub Marketplace recommended, AI Assisted apps/actions, and AI-related GitHub Topics.
  • github_marketplace_categories: 34 GitHub Marketplace category rows.
  • github_scrape_runs: one run record with scrape date, source URLs, row counts, and status.

Why new dataset: your existing datasets are about GoalfyData social/content/outreach, Reddit operations, ecommerce diagnostics, and video workflows. This scrape is a reusable AI-tool discovery/intelligence asset, so a separate dataset is cleaner than mixing it into those.

Reply with confirmed and I’ll create the dataset and import the files.


confirmed

Done. I organized the scrape into a GoalfyData dataset:

GitHub AI Tools Intelligence
Dataset ID: dataset_***
Schema: schema_***

Tables created and loaded:

  • github_ai_tools: 126 rows
  • github_marketplace_categories: 34 rows
  • github_scrape_runs: 1 provenance row

Validation passed: row counts match, key fields have zero nulls, primary keys have no duplicates, and the dataset has registered table metadata, logical relations, governance rules, and a usage guide.

One note: github_ai_tools has 126 source appearances but 119 unique URLs because some repos appear across multiple GitHub Topic surfaces. I preserved those duplicates intentionally so source visibility is analyzable.


Show me the structure of the dataset.

Dataset Structure
GitHub AI Tools Intelligence
ID: dataset_***
Schema: schema_***
Source: github
Mode: manual import at creation; later enabled managed refresh

Tables

github_ai_tools — 126 rows
Primary key: row_id

columntypenullablemeaning
row_idintegernostable row number
scrape_datedatenoscrape date
sourcevarcharnosource bucket, e.g. marketplace_ai_apps
surfacevarcharnonormalized surface
kindvarcharnoapp, action, or repository
namevarcharnotool/listing/repo name
slugvarcharnoMarketplace slug or repo name
descriptiontextyesGitHub description
starsvarcharyesrepo stars, blank for Marketplace rows
languagevarcharyesrepo language, blank for Marketplace rows
urltextnotool/repo URL
source_urltextnopage where row was scraped

github_marketplace_categories — 34 rows
Primary key: category_id

columntypenullablemeaning
category_idintegernostable category row number
scrape_datedatenoscrape date
slugvarcharnoMarketplace category slug
namevarcharnocategory display name
urltextnocategory URL

github_scrape_runs — 1 row
Primary key: run_id

columntypenullablemeaning
run_idintegernoscrape run ID
scrape_datedatenoscrape date
scraped_at_utctimestamptznoscrape completion time
tool_row_countintegernotool row count
category_row_countintegernocategory row count
source_countintegernonumber of source pages
statusvarcharnoscrape status
errors_jsontextyesscrape errors
source_urls_jsontextnosource URL map
local_json_pathtextyeslocal JSON snapshot
local_tools_csv_pathtextyeslocal tools CSV
local_categories_csv_pathtextyeslocal categories CSV
scraper_pathtextyesscraper script path

Relations

  • github_ai_tools.scrape_dategithub_scrape_runs.scrape_date
  • github_marketplace_categories.scrape_dategithub_scrape_runs.scrape_date

Rules

  • Preserve source-surface rows: same repo can appear on multiple GitHub Topic pages.
  • stars and language apply to repository rows, not Marketplace apps/actions.

Help me create a dashboard.

Done. I created and deployed the dashboard app.

App: github-ai-tools-dashboard
App ID: app_***
URL: private app URL hidden

It is private by default, so opening it redirects through GoalfyData login.

Included:

  • Source mix chart
  • Repository language chart
  • Top repositories ranking
  • Marketplace categories
  • Searchable explorer
  • Detail drawer with GitHub links

Verification passed: local API check, production build, GoalfyData preflight, package upload, and deployment status online.


Share this dataset with u***@example.com. Do not share the third table.

Done. I shared GitHub AI Tools Intelligence with u***@example.com.

Access scope:

  • Included: github_ai_tools
  • Included: github_marketplace_categories
  • Excluded: github_scrape_runs

Share ID: ***
Status: pending
Permission: read-only

The invitation email was sent. If they do not see it, ask them to check spam/junk.


Share this dashboard with v***@example.com.

Done. I shared the dashboard with v***@example.com.

Share code: ***
Visibility: specified email only
Permission: read-only
Status: pending
Email sent: yes

If they do not see the invitation, ask them to check spam/junk.


Share this dataset with w***@example.com.

Done. I shared GitHub AI Tools Intelligence with w***@example.com.

Access scope:

  • Included: full dataset

Status: pending
Permission: read-only

The invitation email was sent. If they do not see it, ask them to check spam/junk.


Set it to update automatically every day at 9:00 AM.

I can set the GitHub AI Tools dataset behind the dashboard to refresh automatically every day at 9:00 AM Asia/Shanghai. The dashboard will then show updated data automatically because it reads live from the dataset.

Important: this uses GoalfyData Managed Refresh, and each scheduled run consumes 1 data-update credit. Your current quota had room, but I need your explicit confirmation before enabling a daily schedule.

Please reply confirm daily 9am refresh and I’ll configure it.


confirm daily 9am refresh

Done
Enabled GoalfyData Managed Refresh for GitHub AI Tools Intelligence (dataset_***).

Schedule is 0 9 * * *, timezone Asia/Shanghai, so it updates every day at 9:00 AM.

Verified all three tables through live sandbox runs:

  • github_ai_tools: 126 rows
  • github_marketplace_categories: 34 rows
  • github_scrape_runs: 1 row

Confirmed cron_enabled=true on all three tables.

Updated the dataset description/usage guide so it now says the dataset is scheduled, not manual-only.


Please help me analyze this dataset from GoalfyData: GitHub AI Tools Intelligence.

Let me start by looking up this dataset in GoalfyData.

Found it. Let me now get the full dataset details and table structures.

Good, I have the full dataset structure. Let me now create a task and run comprehensive analysis queries.

Now let me run all the analysis queries in parallel.

Good data. Let me dig deeper with a few more targeted queries.

Excellent data. Now let me compile all findings into a comprehensive analysis report.

Analysis complete. Here's a summary of what the GitHub AI Tools Intelligence dataset reveals:

Dataset Snapshot

  • 126 entries from 6 GitHub surfaces, scraped Aug 5, 2026
  • 119 unique tools — 7 appear across multiple surfaces
  • 3 tables: github_ai_tools, github_marketplace_categories, github_scrape_runs

Market Composition

SegmentCount
Marketplace AI Actions20
Marketplace AI Apps20
Marketplace Recommended Apps26
Open-source repos: ai-agents topic20
Open-source repos: ai-tools topic20
Open-source repos: open-source-ai topic20

Key Findings

  1. Code review dominates AI Marketplace — 15 of 40 AI-specific listings are code review tools, including CodeRabbit, Qodo, CodeAnt, and others. It is the single most crowded AI use case on GitHub.

  2. Claude leads in CI agents — Anthropic has 3 dedicated GitHub Actions versus 1 for Google Gemini CLI, suggesting Claude is the go-to model for AI-in-CI workflows.

  3. Agent frameworks are the hottest trend — The top 4 repositories by stars are all agent-related: ECC, hermes-agent, firecrawl, and langchain.

  4. TypeScript = Python for AI tooling — Both are at 31.7% of open-source AI repos, showing AI has moved beyond Python-only into mainstream developer tooling.

  5. Marketplace AI is still young — Only 40 dedicated AI listings across 34 total Marketplace categories, but 5 AI apps already appear in both “AI Assisted” and “Recommended”, signaling crossover traction.


Governed data asset

Inspect structure, provenance, and refresh rules

A repository may appear on more than one GitHub surface, so source-surface rows are preserved and each update stays traceable to its scrape run.

DatasetGitHub AI Tools Intelligence
Ready

Review core tables, grain, and keys. Hover over another tab to switch views.

github_ai_toolsPreserves the originating GitHub surface
126 rows · repository and Marketplace listings
github_marketplace_categoriesCategory labels and listing counts
34 rows · Marketplace category taxonomy
github_scrape_runsSource URLs, completion time and row counts
1 row · scrape provenance and status
Generated app

Open the live GitHub AI tools dashboard

The public app provides browsable views for source mix, top repositories, language distribution, and Marketplace categories.

Take the result with you

Add “GitHub AI Tools Intelligence” to your account

Enter your email and we’ll send invitations for this dataset and app. Opening an email verifies your address; a new account is created automatically, or you’re signed in to an existing one.

Accepting the dataset uses 1 dataset slot. Your existing account limit still applies.