Skip to main content

How to feed live web data into a RAG pipeline

To feed live web data into a RAG pipeline, crawl the source with Apify's Website Content Crawler, which returns clean Markdown per URL. Chunk that Markdown, embed it, and upsert into a vector store keyed by URL. Re-crawl on a schedule and re-embed only the pages whose content hash changed.

If an answer has to reflect a page as it is right now, skip the index entirely. RAG Web Browser takes a query or a URL and returns cleaned text straight into the prompt.

What Apify does not do: chunking, embedding, and vector storage are yours. Apify owns collect → clean → export. You own embed → store → refresh.

RAG ingestion is one entry in the wider set of Apify use cases — lead lists, price tracking, and SERP data among them — so if web data feeds a different job than a vector store, that page is the better starting point.

What changed since this page was last rewritten​

The page previously recommended an Actor called apify/actors-mcp-server. It does not exist. api.apify.com/v2/acts/apify~actors-mcp-server returns 404 (checked 2026-09-09), because Apify moved MCP from a Store Actor you run to a hosted endpoint you connect to at https://mcp.apify.com. Three links on this page pointed at it. They are gone.

Two other things are newer than the old draft. langchain-apify now ships ApifySearchRetriever, which returns LangChain Document objects directly and removes the loader step for search-shaped retrieval (docs.apify.com/platform/integrations/langchain, checked 2026-09-09). And Apify publishes six official vector-store integration Actors, five of which almost nobody uses — the numbers are in the table below.

Which Actor should you use for RAG data?​

You are ingestingActorWhat comes backWhere it goes
A whole docs site, help centre or blogWebsite Content Crawlermarkdown, text, url, title, metadata per pageChunker, then embeddings
A page you need read at question timeRAG Web BrowserMarkdown, plain text or HTML per requestPrompt context, nothing stored
A known list of URLs, no crawl configFirecrawlLLM-oriented MarkdownChunker, then embeddings

Both Apify Actors carry the FREE pricing model: no rental fee and no per-result charge, so you pay platform usage only (api.apify.com/v2/store, checked 2026-09-09). The Actor page states it plainly: "The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have."

Website Content Crawler offers four crawler modes. Adaptive switching is the default and moves between a headless browser and raw HTTP on its own; Firefox with Playwright handles JavaScript-heavy sites; the raw HTTP Cheerio client is the fastest and cheapest; Chrome with Playwright and JSDOM are marked deprecated (apify.com/apify/website-content-crawler, checked 2026-09-09). Start on adaptive. Drop to raw HTTP once you confirm the pages render server-side, because that single switch is the largest cost lever on this page.

For setup screenshots and input fields, see how to scrape website content with Apify. For a wider field of RAG-adjacent tools, see the best AI data Actors roundup.

What does a RAG crawl actually cost?​

Apify's own estimates, published on the Website Content Crawler page and checked 2026-09-09:

Crawl modeApify's stated cost10,000-page site, one full crawl
Raw HTTP (Cheerio)$0.20 per 1,000 pages~$2
Headless browser$0.50–$5 per 1,000 pages$5–$50
AI summarisation add-on~$2–$3 per 1,000 pages$20–$30 on top

Platform usage is billed in compute units, and the rate falls as the plan grows: $0.20 per CU on Free and Starter, $0.16 on Scale, $0.13 on Business (apify.com/pricing, checked 2026-09-09). The Free plan includes $5 of usage a month; Starter is $19/month subject to change · verified 2026-09-09 — $17 a month billed annually — with $19 of included usage (apify.com/pricing, checked 2026-09-09). See how compute units are counted if the number on your usage chart does not match your expectation.

So a 10,000-page documentation site crawled on raw HTTP fits inside the free monthly credit. The same site on the browser crawler, re-crawled weekly, is $20–$200 a month and is what pushes you onto a paid plan. Crawling is rarely the expensive half of a RAG pipeline. Embeddings and vector hosting are billed by their own vendors and are not included in any figure above.

One conflict to be aware of before you schedule parallel crawls. Apify's two tier-1 sources disagree on Free-plan concurrency: apify.com/pricing lists 5 concurrent Actor runs, while docs.apify.com/platform/limits lists 25 (both checked 2026-09-09). Plan against 5, and confirm on your own account's Limits page before you depend on it.

How do you turn crawled pages into embeddings?​

  1. Crawl. Run Website Content Crawler with your start URLs and a max depth. Use the markdown field as the body text.
  2. Filter. Deduplicate by URL and by body hash, drop pages under a word-count floor, and filter by language. Every page you drop here is an embedding you never pay for.
  3. Chunk. Split on \n\n first, then \n, then sentences. About 1,000 tokens with 10–15% overlap is a reasonable starting point for docs; long manuals do better with parent/child hierarchical chunks.
  4. Embed and upsert. Attach source (the URL), title, and a content hash to every chunk. The source field is what lets the model cite its answer.
  5. Schedule. Put the crawl on an Apify schedule and run the incremental refresh below.

Markdown is the better default for embedding free-form pages, because it keeps headings and lists intact while dropping HTML chrome. Reach for structured JSON only when you need typed fields — price, author, publish date — as filterable metadata on the vector.

Does Apify write to your vector store for you?​

Partly. Apify publishes six official integration Actors, but adoption is concentrated in one of them. User counts and last-modified dates from api.apify.com/v2/acts/..., checked 2026-09-09:

Integration ActorUsersLast modified
apify/pinecone-integration5582026-03-19
apify/qdrant-integration612025-06-09
apify/pgvector-integration162025-04-13
apify/milvus-integration102025-08-15
apify/weaviate-integration72025-04-13
apify/chroma-integration42025-08-25

None is deprecated. But Pinecone is the only one Apify lists on its own integrations page, the only one modified this year, and the only one with adoption in the hundreds. If your vector store is Qdrant, Weaviate, Milvus, Chroma or pgvector, treat the integration Actor as a convenience you should be ready to replace with twenty lines of your own upsert code.

How do you keep the index fresh without re-embedding everything?​

Re-embedding a whole corpus on every run is the default mistake. The fix is a content hash stored beside each vector.

  1. Store a contentHash — SHA-256 of the page Markdown — in each vector's metadata, and use the page URL as the vector id.
  2. Re-run the crawl on your schedule against the same start URLs.
  3. For every row, recompute the hash. Skip the embedding call when it matches. Upsert when it changed or when the URL is new. Delete vectors whose URL no longer appears in the crawl.
import hashlib

def content_hash(s: str) -> str:
return hashlib.sha256(s.encode()).hexdigest()

current_urls = set()
for doc in loader.load():
url = doc.metadata["source"]
current_urls.add(url)
new_hash = content_hash(doc.page_content)
existing = index.fetch(ids=[url]).vectors.get(url)
if existing and existing.metadata.get("contentHash") == new_hash:
continue # unchanged since the last crawl, skip the embedding call
values = embeddings.embed_query(doc.page_content)
index.upsert(vectors=[{
"id": url,
"values": values,
"metadata": {**doc.metadata, "contentHash": new_hash},
}])

How much this saves depends entirely on how often your sources change, and we have not measured that across enough sites to publish a figure. Measure your own: log the number of changed hashes on your first three scheduled runs, and that ratio is your actual saving. A release-notes page will look nothing like a marketing site.

One caveat this pattern hides. Keying vectors by URL means one vector per page, so it only holds while each page fits in a single chunk. Once you split a page into several chunks, key on url#chunk-index and delete every chunk id under a URL before re-upserting it, or stale chunks survive the refresh.

How do you connect Apify to LangChain or LlamaIndex?​

Install langchain-apify and map dataset rows onto Document objects. Package name and class names confirmed at docs.apify.com/platform/integrations/langchain, checked 2026-09-09.

from langchain_apify import ApifyDatasetLoader
from langchain_core.documents import Document
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings
from langchain_pinecone import PineconeVectorStore

loader = ApifyDatasetLoader(
dataset_id="YOUR_DATASET_ID",
dataset_mapping_function=lambda item: Document(
page_content=item.get("markdown") or item.get("text") or "",
metadata={"source": item.get("url"), "title": item.get("title")},
),
)
docs = loader.load()
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=100,
separators=["\n\n", "\n", ". ", " ", ""],
)
chunks = splitter.split_documents(docs)
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = PineconeVectorStore.from_documents(
chunks, embeddings, index_name="YOUR_INDEX_NAME"
)

If you want retrieval without owning a dataset at all, ApifySearchRetriever runs the search and hands back Document objects in one call:

from langchain_apify import ApifySearchRetriever

retriever = ApifySearchRetriever(max_results=3)
docs = retriever.invoke("What is LangChain?")

That is convenient and it is also live retrieval, so it costs a run per query rather than a one-off index build. Use it for low-volume agent lookups, not for a chatbot answering thousands of questions a day. More patterns in the Apify LangChain integration guide.

For LlamaIndex, the reader is ApifyDataset from llama-index-readers-apify, and it takes the same mapping function shape (developers.llamaindex.ai, checked 2026-09-09):

from llama_index.core import Document
from llama_index.readers.apify import ApifyDataset

reader = ApifyDataset(apify_api_token="YOUR_TOKEN")
documents = reader.load_data(
dataset_id="YOUR_DATASET_ID",
dataset_mapping_function=lambda item: Document(
text=item.get("markdown") or item.get("text") or "",
metadata={"source": item.get("url")},
),
)

How do you give an agent live web access instead of a static index?​

Two mechanisms, and they are not interchangeable.

Standby HTTP. RAG Web Browser runs in Apify's Standby mode, so it answers plain GET requests instead of requiring a run to be started and polled:

https://rag-web-browser.apify.actor/search?token=<APIFY_API_TOKEN>&query=hello+world

It searches Google, scrapes the top results in a real browser, and returns Markdown, plain text or HTML. That is a normal HTTP call your agent can make inside a tool handler. See Standby mode for how the warm-instance billing works.

MCP. Apify's Model Context Protocol surface is a hosted endpoint, https://mcp.apify.com, not an Actor you run. Connect over Streamable HTTP with OAuth or a bearer token; for local development, npx @apify/actors-mcp-server speaks stdio. Apify's docs still carry a note that SSE transport "will be removed on April 1, 2026", so build on Streamable HTTP rather than SSE. Setup is documented for Cursor, VS Code and Claude Desktop, and the Apify CLI can write the config for claude-code, cursor, vscode, vscode-insiders, codex, kiro and antigravity (docs.apify.com/platform/integrations/mcp, checked 2026-09-09). Our MCP server setup guide walks through a client config.

Static RAG answers from what you already indexed. MCP and Standby answer from what is on the page this second, and charge you per lookup. Most production systems end up with both: an index for the corpus that changes weekly, live calls for the handful of sources that change hourly.

When is Firecrawl the better tool?​

If you already have the list of URLs and you just want Markdown back, Firecrawl is less work than configuring a crawler. Its free tier is 1,000 credits a month and the cheapest paid tier, Hobby, is $16 a month billed annually for 5,000 credits (firecrawl.dev/pricing, checked 2026-09-09). For a one-off ingest of a few thousand known pages, that is the shorter path and you should take it.

Apify earns its place when the job is a repeating one: full-domain discovery rather than a URL list, scheduling, proxy rotation for sites that block you, and the long tail of awkward sources that need a site-specific Actor rather than a generic crawler. The two also mix cleanly — Apify for the scheduled domain crawl, Firecrawl for the supplementary URLs someone drops in Slack. A fuller breakdown lives in Apify vs Firecrawl.

What to check before you embed​

  • Deduplicate URLs and near-identical bodies.
  • Drop pages under a word-count floor; they are usually errors or placeholders.
  • Confirm the markdown field is populated, not just html.
  • Language-filter if the model serves one locale.
  • Keep source and crawl time on every chunk, for citations and for freshness SLAs.
  • Check robots.txt and the site's terms; prefer public documentation you have the right to ingest. Is Apify legal covers the ground rules.
  • Minimise personal data in anything bound for training or analytics (GDPR, CCPA).

If you need the hosted chatbot layer on top of the corpus rather than building retrieval yourself, compare Chatbase and the alternatives in Chatbase vs Botpress, eesel AI and Intercom before you write any of this code.

Start the crawl​

The realistic first move is small: point Website Content Crawler at one documentation domain on the raw HTTP client, look at the markdown field of twenty rows, and decide whether the extraction is clean enough before you spend anything on embeddings. On Apify's stated $0.20 per 1,000 pages that test costs cents, and the Free plan's $5 monthly credit covers it with room left over. As of 2026, neither Actor charges a rental or per-result fee, so nothing bills beyond platform usage.

Run Website Content Crawler on Apify →

Frequently Asked Questions

Crawl the sources with Apify's Website Content Crawler, which returns clean Markdown per page with navigation and boilerplate stripped. Chunk that Markdown at around 1,000 tokens with 10-15% overlap, embed each chunk, and upsert it into a vector store keyed by URL with the source URL kept in metadata. Then schedule the crawl to re-run and re-embed only pages whose content hash changed. For answers that must reflect a page right now, call RAG Web Browser at query time instead of reading from the index.

Website Content Crawler for anything that looks like a site: docs, help centres, blogs. It crawls a domain and outputs markdown, text, url, title and metadata per page, with nav, headers, footers, ads and cookie banners removed. Add RAG Web Browser when an answer needs a live page rather than an indexed copy. Both carry Apify's FREE pricing model, so there is no rental or per-result fee and you pay only platform usage (checked 2026-09-09).

Apify's own estimates on the Website Content Crawler page are $0.20 per 1,000 pages on the raw HTTP client and $0.50 to $5 per 1,000 pages with a headless browser, with AI summarisation adding roughly $2 to $3 per 1,000 pages (checked 2026-09-09). A 10,000-page docs site is therefore about $2 per full crawl on raw HTTP. Platform usage is $0.20 per compute unit on the Free and Starter plans, dropping to $0.16 on Scale and $0.13 on Business; the Free plan includes $5 of usage a month and Starter is $19 a month. Embeddings and vector hosting are billed separately by their own vendors.

No, not any more. The old apify/actors-mcp-server Actor returns 404 from the Apify API as of 2026-09-09. Apify's MCP surface is now the hosted endpoint https://mcp.apify.com, connected over Streamable HTTP with OAuth or a bearer token, with npx @apify/actors-mcp-server available for local stdio use. Any guide still telling you to add that Actor to your MCP client config is out of date.

Install langchain-apify. ApifyDatasetLoader takes a dataset_id and a dataset_mapping_function that turns each row into a LangChain Document, mapping the markdown or text field to page_content and url and title to metadata. From there you chunk with RecursiveCharacterTextSplitter, embed, and push to a vector store. If you want live retrieval without a stored dataset, ApifySearchRetriever runs the search and returns Documents in one call, though it costs an Actor run per query.

It can, for six stores. Apify publishes integration Actors for Pinecone, Qdrant, pgvector, Milvus, Weaviate and Chroma, none of them deprecated. Adoption is lopsided: Pinecone has 558 users and was last modified in March 2026, while the other five sit between 4 and 61 users and were last touched in 2025 (api.apify.com, checked 2026-09-09). Pinecone is also the only one Apify documents on its own integrations page. For the others, plan to write your own upsert.

If you already have the URLs and just want Markdown back once, yes, Firecrawl is less setup. Its free tier is 1,000 credits a month and Hobby is $16 a month billed annually for 5,000 credits (checked 2026-09-09). Apify is the better fit when the job repeats: full-domain crawling rather than a URL list, scheduling, proxy rotation, and site-specific Actors for sources a generic crawler cannot read. Plenty of pipelines use both.

Markdown for the body text. It keeps headings and list structure while dropping HTML chrome, which produces cleaner chunk boundaries. Use structured JSON when you need typed fields such as price, author, publish date or category as filterable metadata on the vector, for example to restrict retrieval to recent or in-stock items. A common shape is Markdown as the embedded text with a few JSON fields alongside it as metadata.

Match the schedule to how fast the source actually changes, and keep the cost down with a content hash rather than by crawling less often. Store a SHA-256 of each page body in the vector metadata, recompute it on every run, and skip the embedding call when it matches. Documentation and marketing sites usually justify weekly; release notes, status pages and pricing pages justify daily. Log how many hashes change on your first few runs and set the interval from that.

Common mistakes and fixes

Retrieved chunks are navigation and cookie banners, not body text.

Use the `markdown` field, not `html`. Website Content Crawler already strips nav, headers, footers, ads and cookie warnings; if boilerplate still survives, drop pages under a minimum word count before you chunk.

The first sync costs more in embeddings than in crawling.

That is normal. Crawling 10,000 pages on the raw HTTP client is about $2 of Apify platform usage; embedding those same pages is billed by your model provider. Cap crawl depth, dedupe URLs, and embed with a small model until the retrieval quality is proven.

ApifyDatasetLoader returns empty documents.

The mapping function is reading a field the run did not produce. Open the dataset in the Apify console, confirm whether rows carry `markdown` or only `text`, and fall back explicitly: `item.get("markdown") or item.get("text") or ""`.

Your MCP client cannot connect to apify/actors-mcp-server.

That Actor no longer exists. Apify's MCP surface is the hosted endpoint https://mcp.apify.com over Streamable HTTP, or `npx @apify/actors-mcp-server` locally over stdio.