AI & Models · Integration
Replicate
Add Replicate to your product for your customers, and give your AI agents governed access to it.
Replicate runs models as versions, and what you create against a version is a prediction. That prediction is asynchronous by design, and everything else follows from it: you create the prediction, get an identifier back, and then either poll it or let a webhook tell you it finished, because a heavy image or video model can run far longer than any request in your product should stay open. Version pinning matters too, since a version is immutable with a fixed input schema, so a model author publishing something new does not change what your existing calls do until you move to it. Outputs arrive as files to fetch rather than inline data. Tokens are held per customer by fastn, which also keeps the prediction endpoints current.
In your product
Embedded for your customers. Per-tenant auth, no per-customer code, maintained by fastn.
Let each customer connect their own Replicate account so generation costs sit on their token instead of yours.
Create a prediction and hand control straight back to the user, finishing the work when the webhook lands.
Pin the model version your product calls, so a newly published version cannot quietly change your output overnight.
Fetch and store the output files a prediction produces, rather than expecting a result inside the response body.
For your AI agents
Governed, audited access for the agents you build, through the MCP server.
An agent starts a prediction, reports the identifier, then answers once the result is ready.
An agent reads a version's input schema before assembling the arguments for a run.
An agent lists a customer's recent predictions with their state and how long each one took.
Example prompt
Is this prediction still running, and which model version did I start it on?
Set up Replicate in 4 steps
- 01Enable the Replicate connector from your fastn dashboard.
- 02Have each customer authorise their own Replicate account, so calls run under their credentials rather than a shared key.
- 03Decide which model versions and predictions your product needs, and plan for predictions being asynchronous, so you create one then poll it or take a webhook.
- 04Call it from your product and expose it to your agents through the same governed connection.
Why teams use the Replicate integration
What you get by embedding it with fastn instead of building it yourself.
- Ship a Replicate integration without building it. Your customers connect their own Replicate account inside your product and work their models, requests and outputs there, with no per-customer code on your side.
- Handle the part that actually costs time: per-customer keys, quotas and cost attribution matter more than schema here, because every call is billed. fastn owns the auth, token refresh, rate limits, pagination and breaking-change fixes, so a Replicate update is not your on-call problem.
- One integration serves your product and your agents. The same governed Replicate connection powers in-product features and gives AI agents scoped, audited access, so you give your product and your agents a model call each customer pays for themselves without wiring it twice.
Used by these teams
Compare with
Often used alongside
Tools the same teams tend to run next to Replicate, across other categories.
Replicate integration FAQ
How do I add a Replicate integration to my product?
Enable the Replicate connector in your fastn dashboard, then let each customer authenticate their own Replicate account. fastn handles the OAuth flow, token storage and refresh per tenant, so there is no Replicate client code in your app and no per-customer branch in your codebase. Setup is 4 steps.
Do my customers each connect their own Replicate account?
Yes. Every connection is scoped to the individual customer, so each authorises their own Replicate account and only ever sees their own models, requests and outputs. That per-tenant isolation is the point of an embedded integration: you support the long tail of customer setups without maintaining an integration per customer.
Can AI agents use this Replicate integration?
Yes. The same connection is exposed to your agents through the fastn MCP gateway, with permissions scoped per tenant and every call audited. An agent starts a prediction, reports the identifier, then answers once the result is ready.
Who maintains the Replicate integration?
fastn does. When Replicate changes an endpoint, deprecates a field or alters its auth, the fix lands in the connector rather than in your backlog, and your customers' connections keep working.
Whose Replicate API key and quota does each call use?
Each customer authorises their own Replicate account, so usage, rate limits and cost land on the customer that caused them. You are not metering a shared key and re-billing it, and one heavy customer cannot exhaust another's quota.
Are inputs and outputs auditable?
Yes. Every call is logged per tenant with the call, the input and the result, so an output can be traced back to what produced it. That matters more here than in most integrations, because a generated answer or an extracted field cannot be reconstructed from the request alone.
What can I build with the Replicate integration?
A common starting point: let each customer connect their own Replicate account so generation costs sit on their token instead of yours. Teams also use it for the other use cases listed above, and expose it to agents for governed reads and writes.
How much does the Replicate integration cost?
It is included. Pricing is based on connected accounts, not on how many connectors you enable, so adding Replicate does not change your per-connector cost. You can start free with 3 connected accounts.
Add Replicate to your product
Start free with 3 connected accounts. No sales call required, and no per-customer integration code.