Connect YouTube to Chatbase: Timestamped Video Answers
Chatbase YouTube Connector is a free, MIT-licensed tool that keeps a Chatbase AI agent trained on your YouTube channel, so the agent answers customers from your videos and links to the exact minute it used. It runs as a scheduled GitHub Action or a Docker container on any VPS, with no server or database of its own. Transcripts come from an Apify Actor at about $1 per 1,000 videos, and a day with no new uploads costs $0. On Chatbase Standard and above it syncs through the API; on Free and Hobby it exports ready-to-upload files instead.
Here is what a customer sees once it is running:
Customer: How do I hand a chat over to a human? Agent: Turn on the Chatbase live chat action, then ... watch at 0:24
Why can't Chatbase train on YouTube videos today?
Chatbase can train an agent on files, websites, text, Q&A, Notion and tickets. YouTube is not on that list, so a support team with a library of product walkthroughs has two options: download and paste transcripts by hand, or leave that knowledge out of the agent. Neither keeps up when you publish every week.
The connector fills that gap. It finds new videos, fetches their transcripts, formats them with timestamps and chapter headings, and writes each video into Chatbase as a text source. It also handles the parts that are easy to get wrong in a one-off script: not paying twice for the same video, not duplicating sources, and not deleting half your knowledge base by accident.
What does a timestamped answer look like?
Every video becomes one Chatbase source, split into sections of about one minute. Each section ends on a sentence and carries a ?t= link to that moment. When the video description has chapters, they become the section headings.
This is a trimmed excerpt of a real source the connector produced from a video on Chatbase's own channel:
# Chatbase Backstage - Manage Your AI Agent with Natural Language (video)
Channel: Chatbase · Published: 2026-07-22 · Duration: 2:53 · Language: en · Transcript: auto-generated captions
URL: https://www.youtube.com/watch?v=fq3x-zpy5JY
When you answer from this video, link the timestamp of the section you used.
## What Backstage is · 0:00–0:45 · [watch](https://youtu.be/fq3x-zpy5JY?t=0)
Hello, my name is Sia, and today I'm going to walk you through backstage in Chat base. [...]
## As an analyst: reports, blind spots, and conflicts · 0:43–1:27 · [watch](https://youtu.be/fq3x-zpy5JY?t=43)
In the agent menu, navigate to backstage. First, let's type, "Analyze my last 100 conversations. [...]
## As an operator: building an action · 1:26–2:06 · [watch](https://youtu.be/fq3x-zpy5JY?t=86)
Now, let's have a look at the actions it can take. Say, I run an online store and I want to capture interested visitors as leads. [...]
The header tells the agent what the source is, and the line "When you answer from this video, link the timestamp of the section you used" travels with the content. The transcript text is the captions as YouTube has them, so auto-generated captions keep their quirks ("Chat base" above).
How does the connector work?
The pipeline has five steps:
YouTube RSS / channel listing ─► Apify transcript Actor ─► format (timestamped Markdown) ─► diff vs Chatbase ─► create / update / delete text sources
There are two kinds of run:
| Run | Discovery | What it does |
|---|---|---|
sync (daily) | Free: channel RSS (newest 15), playlists in full | Transcribes and adds videos you don't have yet |
sync --full (weekly) | Every upload (including Shorts and live recordings) via the Actor | Also updates changed transcripts and, with prune: true, removes videos deleted from YouTube |
The daily run is cheap because discovery is free: YouTube's RSS feed for channels and the playlist page for playlists. Videos that have no captions or are too short to keep go into a small skip cache, so the connector does not pay to check them again every day. It re-checks skipped videos after 30 days by default, in case captions were added.
State lives in Chatbase, not in a database. Each source is named YT·<videoId>·<hash>·<title>. On every run the connector reads the agent's existing source names, compares them with what YouTube has, and only writes the difference. A re-run with nothing new makes zero writes, which is why the runner can be thrown away after each job. The skip cache is only an optimisation: losing it means skipped videos get checked once more.
Transcripts come from the YouTube Transcript Scraper Pro Actor on Apify. It returns timestamped captions and, if you enable it, AI speech-to-text for videos that have no captions.
How do you set it up with GitHub Actions in 5 minutes?
This route needs a Chatbase Standard plan or higher, because it writes through the Chatbase API. (On Free or Hobby, skip to export mode.)
1. Add a config file. Create chatbase-youtube.yaml in any repository:
version: 1
jobs:
- name: academy
agentId: ${CHATBASE_AGENT_ID}
sources:
- channel: '@YourChannel'
2. Add secrets under Settings → Secrets and variables → Actions:
| Name | Type | Where to get it |
|---|---|---|
APIFY_TOKEN | Secret | Create a free Apify account, then Settings → API & Integrations |
CHATBASE_API_KEY | Secret | Chatbase → Workspace settings → API keys (Standard plan or higher) |
CHATBASE_AGENT_ID | Variable | The ID in your agent's URL |
3. Copy the workflow. Save the repo's examples/workflows/daily-sync.yml as .github/workflows/chatbase-youtube.yml. The core of it:
on:
schedule:
- cron: '17 6 * * *' # daily, new videos only (free RSS discovery)
- cron: '43 4 * * 0' # Sunday, full re-check
workflow_dispatch:
inputs:
full: { type: boolean, default: false, description: Full re-check }
dry-run: { type: boolean, default: false, description: Plan only }
jobs:
sync:
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- uses: actions/checkout@v4
- uses: actions/cache@v4
with:
path: .chatbase-youtube
key: chatbase-youtube-state-${{ github.run_id }}
restore-keys: chatbase-youtube-state-
- uses: yel-hadd/chatbase-youtube-connector@v1
with:
config: chatbase-youtube.yaml
full: ${{ github.event.schedule == '43 4 * * 0' || inputs.full == true }}
dry-run: ${{ inputs.dry-run == true }}
env:
APIFY_TOKEN: ${{ secrets.APIFY_TOKEN }}
CHATBASE_API_KEY: ${{ secrets.CHATBASE_API_KEY }}
CHATBASE_AGENT_ID: ${{ vars.CHATBASE_AGENT_ID }}
The full example file also sets a concurrency group so two runs never overlap, and uploads report.json as an artifact after every run. The cache step keeps the skip cache between runs, and because it refreshes daily, GitHub does not evict it.
4. Dry run first. Open the Actions tab, run the workflow manually with dry-run ticked, and read the plan on the run page. A dry run spends nothing on Apify and writes nothing to Chatbase. When the plan looks right, run it again without the box ticked. After that, the schedule takes over: new videos daily, a full re-check every Sunday.
Each run writes report.json with counts, every video's outcome, estimated and actual spend, and storage before and after. The same summary appears on the GitHub run page.
How do you run it on a VPS with Docker?
If you would rather keep it off GitHub, the Docker image has its own scheduler:
mkdir -p chatbase-youtube/data && cd chatbase-youtube
sudo chown 1000:1000 data # the container runs as an unprivileged user (uid 1000)
curl -O https://raw.githubusercontent.com/yel-hadd/chatbase-youtube-connector/main/docker-compose.yml
curl -o chatbase-youtube.yaml https://raw.githubusercontent.com/yel-hadd/chatbase-youtube-connector/main/examples/chatbase-youtube.yaml
printf 'APIFY_TOKEN=...\nCHATBASE_API_KEY=...\nCHATBASE_AGENT_ID=...\n' > .env && chmod 600 .env
docker compose run --rm chatbase-youtube-sync doctor # check everything, spend nothing
docker compose run --rm chatbase-youtube-sync sync --dry-run # see the plan
docker compose up -d # daily sync + weekly full re-check
Do not skip the chown line. The container runs as a non-root user (uid 1000) with a read-only filesystem, and ./data is the one writable mount, where it keeps the skip cache, report.json and export-mode files. If the host directory belongs to root, the container cannot write there.
doctor checks your tokens, Chatbase plan access, the agent and the channels without spending anything. The compose file schedules a daily sync and a full re-check on Sundays; change them with the SCHEDULE and FULL_SCHEDULE environment variables in cron syntax. Two optional variables help with monitoring: HEALTHCHECK_URL is pinged on start, success and failure (healthchecks.io style), and REPORT_WEBHOOK_URL receives report.json after each run. If you prefer systemd to the built-in scheduler, the repo ships unit files in deploy/.
Chatbase Free or Hobby: use export mode
The Chatbase API starts at the Standard plan ($150/month, or $120/month billed annually). Free and Hobby workspaces cannot use it, so the connector has a second output: set sink: export in the job.
In export mode it writes one .txt file per video plus a CHANGES.txt that lists exactly which files to upload or delete under Sources → Files in the Chatbase dashboard. Re-runs only rebuild changed videos, so after the first batch you only handle what is new. Only APIFY_TOKEN is needed.
The repo's examples/workflows/export-free-plan.yml runs weekly (Mondays) and commits the new and changed files back to your repository. Committing, rather than caching, keeps the state safe between weekly runs. Each week you open CHANGES.txt and upload the files it lists.
Mind the storage limit on the smaller plans: Free allows 1 MB of training content and Hobby 10 MB. See our Chatbase pricing breakdown for what each plan includes.
How do you choose which videos go in?
sources is an allow-list: only what you list is synced. Mix channels, playlists and single videos, then carve out what you don't want:
jobs:
- name: academy
agentId: ${CHATBASE_AGENT_ID}
sources: # only what you list is synced
- channel: '@AcmeAcademy' # a whole channel...
- playlist: 'PLxxxxxxxx' # ...or just some playlists
- video: 'https://youtu.be/xxxxxxxxxxx' # ...or single videos
exclude: # keep these out, even if a source includes them
- 'https://youtu.be/yyyyyyyyyyy'
- 'PLzzzzzzzz' # every video in this playlist
filters:
titleExclude: ['(?i)teaser|trailer']
publishedAfter: '2024-01-01'
prune: true # also remove excluded or deleted videos already in the agent
A few rules worth knowing:
- Excluded videos cost nothing. They are never transcribed. If one is already in the agent, the run report flags it, and with
prune: truethe next run removes it. - Filters apply to new videos only.
titleInclude,titleExcludeandpublishedAfternever delete a video that is already in the agent. Onlyexcludeplusprune, or a deletion on YouTube, removes one. - Shorts and very short videos are skipped by default.
includeShortsisfalseandminDurationSecis60. - Big playlists need a YouTube API key. Without
YOUTUBE_API_KEY, the connector reads the public playlist page, which shows the first 100 videos. For larger playlists the run warns you; a free YouTube Data API v3 key gives a complete listing. - Other languages.
languagessets caption preference (default[en]), andmachineTranslate(on by default) translates captions from another language into your first one while keeping timestamps.
The full option list is in the repo's docs/configuration.md. One config file can hold several jobs, one per Chatbase agent.
What does it cost? A worked example
Apify is the only cost the connector adds on top of your Chatbase plan. The connector's estimate uses $0.001 per transcript and $0.012 per minute of AI speech-to-text; both prices are configurable under pricing.
| Scenario | Apify cost |
|---|---|
| Daily run, no new videos | $0 (free discovery, skip cache) |
| 1 new video with captions | about $0.001 |
| Backfill 300 videos × 20 min, 20% without captions (AI) | 300 × $0.001 + 60 × 20 × $0.012 ≈ $14.70 |
The backfill line shows where the money goes: the 300 caption transcripts cost $0.30, and the 60 videos that need AI transcription cost $14.40. AI fallback is off by default (aiFallback.enabled: false). If you turn it on, aiFallback.maxMinutesPerRun (default 60) and aiFallback.skipLongerThanMin (default 90) cap it.
Chatbase's own cost is the plan you are already on. The connector writes sources; it does not send chat messages, so it uses no message credits. Storage is the limit to watch: an hour of speech is roughly 55 to 65 KB of text, and the plans allow 1, 10, 20 or 40 MB of training content (Free, Hobby, Standard, Pro). A 20 MB Standard plan therefore holds roughly 300 hours of video transcripts if nothing else is in the agent.
How do you make the agent cite timestamps?
The timestamps only help if the agent passes them on. Add this to your agent's instructions in Chatbase:
When an answer comes from a video source (titles ending in "(video)"), include the matching
[watch](…)link from that section so the customer can jump to the moment. If no source covers the question, reply exactly: "I couldn't find this in our videos."
The second sentence matters as much as the first. It stops the agent from filling gaps with guesses when your videos don't cover a question.
How safe is it to run unattended?
The defaults are built so a misconfigured run fails closed:
- Budget guard. Before any Apify run, the connector estimates the worst-case spend. If it exceeds
budget.maxUsdPerRun(default $5), the run aborts before spending anything, with exit code 3.budget.maxNewVideosPerRun(default 200) caps new videos per incremental run; the rest wait for the next run. - Dry run and doctor.
--dry-runshows the plan without spending or writing.doctorchecks tokens, plan access, the agent and channels. - Delete cap. Deletions are opt-in (
prune: falseby default) and capped atbudget.maxDeletesPerRun(default 10). Anything above that is held back unless you pass--allow-mass-delete. A full run that hit themaxVideoslisting limit (default 500) skips pruning, because a truncated list could make real videos look deleted. - Storage limit. Set
budget.storageLimitMbto your plan's training limit. The run adds videos until the next one would go over, then stops with exit code 3. - Rate limits. It respects Chatbase's limit of 100 requests per 10 seconds and handles
409and429responses. - Secrets and data. Tokens come only from environment variables and are redacted in logs. There is no telemetry: data moves only between YouTube, Apify and your Chatbase workspace. Use it for your own channel or content you have rights to; only public and unlisted videos are processed.
Exit codes make failures easy to alert on: 2 is an invalid config, 3 a budget, storage or delete cap, 4 a partial failure where some videos failed and the rest synced, and 5 an auth or plan problem such as the API needing Standard.
How we built and tested this
We built the connector at use-apify.com. Version 0.1.0 was tested live on Chatbase's own YouTube channel, @chatbase_, in export mode; the transcript excerpt above is real output from that run. The Chatbase API sink was written and checked against Chatbase's published OpenAPI specification for the v2 Sources API, but we have not yet run it against a live Standard workspace. That test is pending Standard plan access, and we will update this post when it is done. Until then, export mode is the path we have run on real videos, and a dry run is the way to check the API plan against your own agent before anything is written.
If you are new to the API side, our Chatbase API integration guide covers keys, rate limits and the Sources API this tool writes to. For a wider view of the platform, read our Chatbase review or the longer Chatbase review of the AI chatbot builder.
To try it, create a free Chatbase agent and start with export mode, or compare Chatbase plans if you want automatic syncing through the API. The code is on GitHub.
FAQ
Not natively. Chatbase trains on files, websites, text, Q&A, Notion and tickets, but has no YouTube source. The open-source Chatbase YouTube Connector fills the gap by turning each video's transcript into a Chatbase text source with a timestamped link about every minute.
Any plan works. On Standard and above ($150/month, or $120/month billed annually) the connector syncs automatically through the Chatbase API. On Free and Hobby, which have no API access, set sink: export and it writes one .txt file per video plus a CHANGES.txt listing what to upload under Sources, Files.
Transcripts come from an Apify Actor at about $0.001 per video, so roughly $1 per 1,000 videos. A daily run with no new uploads costs $0. Optional AI speech-to-text for videos without captions costs about $0.012 per minute. Backfilling 300 twenty-minute videos where 20% need AI transcription comes to about $14.70.
No. The connector only creates, updates and deletes text sources; it does not send chat messages. The Chatbase limit to watch is training storage: about 55 to 65 KB of text per hour of speech, against 1, 10, 20 or 40 MB on Free, Hobby, Standard and Pro. Set budget.storageLimitMb so a run stops before going over.
By default they are skipped and remembered in a skip cache, so they are not paid for again every day; they are re-checked after 30 days in case captions were added. If you enable aiFallback, the connector transcribes them with AI speech-to-text, capped by aiFallback.maxMinutesPerRun and aiFallback.skipLongerThanMin.
No. It only manages sources whose names start with YT· and the video ID, and a re-run with nothing new makes zero writes. Deletions are off unless you set prune: true, and even then they are capped at 10 per run by default unless you pass --allow-mass-delete.
No. The easiest setup is a scheduled GitHub Action: add a config file, two secrets and a variable, and copy the example workflow. If you prefer your own machine, a Docker image runs on any VPS with a built-in daily and weekly schedule, or you can use the provided systemd units.
Yes. sources is an allow-list of channels, playlists and single videos. exclude removes specific videos or whole playlists, and filters can include or drop titles by regex or skip videos published before a date. Shorts and videos under 60 seconds are skipped by default.
