What is an Apify Actor?
An Apify Actor is a program packaged as a Docker image that Apify runs on its own machines. It reads one JSON input object, does the work, and writes rows to a dataset plus optional files to a key-value store. You start it from the Console, the API, the CLI, or a schedule.
Apify's own wording: "serverless cloud programs that take a structured JSON input, perform a task (web scraping, browser automation, data processing, and more), and optionally produce a structured output" (Actors documentation, fetched 2026-09-09).
What an Actor is not: it is not an API you get to design. The interface is fixed — JSON in, dataset rows out — and every Actor in the Store obeys it. That constraint is the product. It is why you can replace one Google Maps scraper with another and change nothing but the input form.
Every number below carries the date it was fetched from an Apify page. Prices move quarterly. The structure has not changed in years.
Why does an Actor have this shape?
Four properties fall out of the container-plus-fixed-IO design, and they are the reason the Store works at all:
- Reproducible. A run stores the build it used and the input JSON it was handed, so "what exactly ran on Tuesday?" has an answer on Friday.
- Scalable. The platform places runs; you set memory and let it start many in parallel, up to your plan's ceiling.
- Composable. One Actor can start another or hand results off by webhook to n8n or Make.
- Portable. If it runs in a Linux container, it runs here. JavaScript and Python get first-party SDKs; anything else talks to the HTTP API.
You pay for that with real constraints. There is no long-lived process, no local disk you can rely on between runs, and no interactive session — state has to live in a dataset, a key-value record, or a request queue, or it does not exist after the container exits. Apify published the reasoning as the Web Actor programming model whitepaper.
What happens during a run?
- Bootstrap. Apify starts the container and injects the environment: token, default storage IDs, proxy settings if you configured them.
- Input. Your code loads one JSON object, whether the run came from the Console form, the API, the CLI, or a scheduled task.
- Work. Crawling, browser automation, API calls, transforms — whatever the image does.
- Output. Rows go to the default dataset; screenshots, HTML dumps and checkpoints go to the key-value store.
- Exit. A clean return marks the run succeeded.
Actor.fail('Could not finish the crawl, try increasing memory')marks it failed with a status message you can read in the run list (basic commands, fetched 2026-09-09).
How does an Actor read its input?
At runtime it is one JSON object, read through the SDK.
import { Actor } from 'apify';
await Actor.init();
const input = (await Actor.getInput()) ?? {};
const startUrl = input.startUrl ?? 'https://example.com';
const maxPages = input.maxPages ?? 10;
Python uses a context manager instead of the explicit init/exit pair:
import asyncio
from apify import Actor
async def main():
async with Actor:
actor_input = await Actor.get_input() or {}
start_url = actor_input.get('startUrl', 'https://example.com')
if __name__ == '__main__':
asyncio.run(main())
The declared schema: .actor/INPUT_SCHEMA.json
The runtime object is only half the contract. Publishing Actors ship .actor/INPUT_SCHEMA.json, and that file is what renders the Console form, validates types before a run starts, and documents defaults for anyone integrating over the API. Only schemaVersion 1 exists, and the file is capped at 500 KB (input schema specification, fetched 2026-09-09).
{
"title": "Example scraper input",
"type": "object",
"schemaVersion": 1,
"properties": {
"startUrl": {
"title": "Start URL",
"type": "string",
"description": "First page to crawl",
"editor": "textfield"
},
"maxPages": { "title": "Max pages", "type": "integer", "default": 10 }
},
"required": ["startUrl"]
}
Write this file early. Change a field name after people have automated against it and you have broken their scheduled runs, not just your form.
Where does the output go?
Three storages, three different jobs:
| Storage | Best for | Written with |
|---|---|---|
| Dataset | Many similar records: products, posts, map pins | Actor.pushData(record) |
| Key-value store | Files and single blobs: screenshots, raw HTML, crawl state | Actor.setValue(key, value, { contentType }) |
| Request queue | The crawler frontier — URLs plus metadata | Crawlee and SDK queue APIs |
setValue defaults to JSON, so binary payloads need the content type spelled out: await Actor.setValue('screenshot.png', buffer, { contentType: 'image/png' }). Read it back with Actor.getValue(key) in a later run to resume work.
On the consuming side, results come out of the API rather than a download button: GET https://api.apify.com/v2/datasets/:datasetId/items returns json, jsonl, xml, html, csv, xlsx or rss, and needs a token (dataset items endpoint, fetched 2026-09-09). The Apify API tutorial has the auth and pagination detail; the storage guide covers retention.
A complete Actor, end to end
Input you pass in the Console or API:
{
"startUrls": [{ "url": "https://news.ycombinator.com" }],
"maxItems": 5
}
src/main.js:
import { Actor } from 'apify';
import { CheerioCrawler } from 'crawlee';
await Actor.init();
const input = await Actor.getInput();
const startUrls = input?.startUrls ?? [{ url: 'https://news.ycombinator.com' }];
const maxItems = input?.maxItems ?? 5;
let count = 0;
const crawler = new CheerioCrawler({
maxRequestsPerCrawl: maxItems,
async requestHandler({ request, $ }) {
await Actor.pushData({
url: request.url,
pageTitle: $('title').text().trim(),
scrapedAt: new Date().toISOString(),
});
count += 1;
},
});
await crawler.run(startUrls);
await Actor.setValue('OUTPUT', { itemCount: count });
await Actor.exit();
One dataset row out:
{
"url": "https://news.ycombinator.com",
"pageTitle": "Hacker News",
"scrapedAt": "2026-09-09T12:00:00.000Z"
}
That is the whole model. The rest is scale, blocking, and billing.
What does running an Actor cost?
Two meters run at once, and confusing them is how people get surprised by an invoice.
Meter one is platform usage, measured in compute units: memory multiplied by runtime. Your plan comes with a monthly credit that this is billed against.
| Plan | Monthly | Annual, per month | Included usage | Price per CU | Max RAM |
|---|---|---|---|---|---|
| Free | $0 | $0 | $5 | — | 16 GB |
| Starter | $19 | $17 | $19 | $0.20 | 64 GB |
| Scale | $199 | $179 | $199 | $0.16 | 256 GB |
| Business | $999 | $899 | $999 | $0.13 | 512 GB |
Source: apify.com/pricing, fetched 2026-09-09. Starter is $19/month subject to change · verified 2026-09-09. The included usage is prepaid credit, not an access fee — and Apify states plainly that unused credit does not roll over and expires at the end of the billing cycle. You only pay more when you exceed it. See pricing and cost optimization for the arithmetic.
Meter two is what the Actor's author charges you, and it changed shape this year. The models Apify documents today are pay-per-event, where the code calls a charge for each unit it delivers, and pay-per-usage, where you pay nothing beyond platform costs (monetization docs, fetched 2026-09-09). The old flat monthly rental is being retired on a published schedule: no new rental Actors or rental price changes since 1 April 2026, and on 1 October 2026 the remaining ones are migrated to pay-per-usage (rental docs, fetched 2026-09-09).
The migration is visible in the wild. apify/instagram-scraper billed per dataset item until 12 December 2025 and has been pay-per-event since. bebity/linkedin-jobs-scraper is still on a $29.99 flat monthly price today and flips to pay-per-event on 15 September 2026 (both checked against api.apify.com/v2/acts/... on 2026-09-09). Any monthly figure you read in an older tutorial is about to be wrong.
Pay-per-event prices are also tiered by plan, which most write-ups miss. On compass/crawler-google-places the place-scraped event costs $0.004 on the free tier and $0.003 on Bronze, the tier a Starter subscription puts you in (live API response, 2026-09-09). Deeper detail lives in pay-per-event pricing and, if you are publishing rather than buying, monetize your Actors.
Check what an Actor charges, before you run it
The Store page shows the price, but the API shows the schedule:
curl -s "https://api.apify.com/v2/acts/bebity~linkedin-jobs-scraper" \
| jq '.data.pricingInfos[-1] | {pricingModel, startedAt}'
Read the result carefully. pricingInfos is a historical array, and future changes sit at the end of it alongside past ones. On 2026-09-09 that command returns PAY_PER_EVENT with a startedAt of 2026-09-15 — a model that has not started yet. The price you will actually be charged today is the previous entry, $29.99 a month. Take the last entry whose startedAt is in the past, and compare against a full timestamp rather than a bare date, or a change made earlier the same day reads as still upcoming.
Which limits do you trust when Apify's own pages disagree?
Design against the free plan's concurrency and you hit a genuine contradiction between two first-party pages:
| Limit | apify.com/pricing | docs.apify.com/platform/limits |
|---|---|---|
| Free plan, max concurrent Actor runs | 5 | 25 |
| Free plan, max memory | 16 GB | 16,384 MB |
| Starter, max concurrent Actor runs | 32 | 32 |
Both pages fetched 2026-09-09. They agree on memory and on every paid tier; the free-plan concurrency numbers are five times apart and Apify has not reconciled them. Assume 5 until a run proves otherwise. A fan-out designed for 25 that gets 5 does not fail loudly — it queues, and your scheduled job quietly takes five times as long.
How many Actors are in the Store, really?
Apify's Store headline reads "70,000+ ready-to-run tools called Actors" (apify.com/store, fetched 2026-09-09). The public Store API returns a different figure the same day: total of 57,363 from api.apify.com/v2/store.
Neither is wrong. They count different things — the API total covers the public, searchable listing, and the headline does not say what it includes. Quote "70,000+" if you are quoting Apify. Use the API number if you need one you can reproduce. Do not average them.
Whichever number you take, the practical problem is unchanged: search returns dozens of Actors for the same site, at different prices and quality. That is what our best Apify Actors category pages exist to narrow, and the Store guide covers how to read a listing.
When is an Actor the wrong tool?
Plenty of the time, and the platform is a poor fit in four recognisable cases.
A one-off job on a site that does not block you. Crawlee, the Apache-2.0 library Apify itself maintains and most Actors are built on, runs on your laptop with no account and no bill. If you need 2,000 pages once, from a site that serves plain HTML, deploying it to a cloud runtime buys you nothing.
Data that must not leave your infrastructure. Runs happen on Apify's machines and results sit in Apify's storage. For regulated data or an internal system behind SSO, self-host Crawlee instead.
You only want IP addresses. If your scraper is written, working, and just needs clean residential exits, buy proxies from a proxy vendor. Apify Proxy is competitive inside the platform, but wrapping working code in an Actor to get IPs adds a runtime you do not need.
Production work on the free plan. $5 of monthly credit that expires and a concurrency ceiling of 5 is a test budget. It is enough to learn the input contract on a small run, not to schedule anything you depend on. The free plan guide covers what actually fits inside it.
How do I run an Actor without writing code?
- Open an Actor in the Apify Store — for example the Google Maps or Instagram scrapers.
- Read the Pricing tab first, and note whether it charges per event or per month.
- Fill the input form, and set the results limit to something small like 10.
- Click Start and watch the log until the run finishes.
- Open Storage → Dataset and export JSON, CSV or Excel, or copy the dataset ID for API access.
Repeat it on a schedule once the input is right.
Run one Store Actor with a 10-result limit. Common objection: you do not want to spend anything to find out. You largely do not — pay-per-event Actors bill per unit delivered, so 10 places from compass/crawler-google-places costs $0.04 on the free tier at the price verified above, inside the $5 monthly credit. That is still the fastest way to understand inputs and outputs in 2026. Open the Store.
An Apify Actor is a program packaged as a Docker image and run on Apify's cloud. It reads a single JSON input object, performs a task such as scraping or browser automation, and writes structured rows to a dataset plus optional files to a key-value store. Apify's own docs define Actors as serverless cloud programs that take structured JSON input and optionally produce structured output (fetched 2026-09-09).
Two charges apply. Platform usage is billed in compute units against your plan's monthly credit: the Free plan includes $5, Starter is $19/month with $19 of included usage at $0.20 per CU (apify.com/pricing, fetched 2026-09-09). On top of that, many Store Actors add author charges, now normally pay-per-event. Check an Actor's Pricing tab before a large run.
Yes, on a published schedule. Since 1 April 2026 no new rental Actors can be published and rental prices cannot be changed, and on 1 October 2026 remaining rental Actors are migrated to pay-per-usage pricing (docs.apify.com rental documentation, fetched 2026-09-09). Any monthly rental price quoted in an older tutorial is close to its expiry date.
Apify's own pages disagree. apify.com/pricing says 5 concurrent Actor runs on the Free plan; docs.apify.com/platform/limits says 25. Both were fetched on 2026-09-09 and neither has been corrected. Design for 5 — the failure mode of being wrong is silent queueing, not an error.
Crawlee is the Apache-2.0 open-source crawling library Apify maintains, and it runs anywhere, including on your own machine with no Apify account. An Actor is the cloud package that runs on Apify — often built with Crawlee — plus the platform's input schema, storage, scheduling, API, and Store distribution.
No. Store Actors ship an input form generated from their .actor/INPUT_SCHEMA.json file, so you fill in fields, click Start, and export the dataset as JSON, CSV or Excel. Code is only required when you build your own Actor or wire runs into another system through the API.