AI & Models · Integration
Hugging Face
Add Hugging Face to your product for your customers, and give your AI agents governed access to it.
Hugging Face is a hub and a hosting platform at once, so a Hugging Face integration touches hosted models you call for inference, the dedicated inference endpoints a customer deploys for themselves, and the model repositories, datasets and spaces their token is allowed to see. The behaviour to design around is the cold start. An endpoint that has been idle can be scaled down to nothing, and the first request after that pays for the model to load, which is seconds rather than milliseconds and is not a failure. Anything user-facing therefore needs a timeout that expects it and something to show while it happens. Private repositories add a second axis, since visibility differs per customer. fastn holds each account's token per tenant and follows changes on the inference API.
In your product
Embedded for your customers. Per-tenant auth, no per-customer code, maintained by fastn.
Let each customer bring their own Hugging Face token so inference runs on their account and against their quota.
Call a customer's dedicated inference endpoint from your product, treating the first slow request as a normal path rather than an error.
List during onboarding which model repositories and datasets a customer's token can genuinely reach, private ones included.
Let a customer point your product at a specific model revision or space instead of a default you picked for them.
For your AI agents
Governed, audited access for the agents you build, through the MCP server.
An agent runs inference on the model a customer selected and cites which endpoint answered it.
An agent reads a model card before it recommends that model for a job.
An agent checks whether an endpoint is already warm before promising a fast answer.
Example prompt
Is this inference endpoint warm right now, and which model revision is it serving?
Set up Hugging Face in 4 steps
- 01Enable the Hugging Face connector from your fastn dashboard.
- 02Have each customer authorise their own Hugging Face account, so calls run under their credentials rather than a shared key.
- 03Decide which models, requests and outputs your product needs, map those fields, then enable the actions and triggers you want.
- 04Call it from your product and expose it to your agents through the same governed connection.
Why teams use the Hugging Face integration
What you get by embedding it with fastn instead of building it yourself.
- Ship a Hugging Face integration without building it. Your customers connect their own Hugging Face account inside your product and work their models, requests and outputs there, with no per-customer code on your side.
- Handle the part that actually costs time: per-customer keys, quotas and cost attribution matter more than schema here, because every call is billed. fastn owns the auth, token refresh, rate limits, pagination and breaking-change fixes, so a Hugging Face update is not your on-call problem.
- One integration serves your product and your agents. The same governed Hugging Face connection powers in-product features and gives AI agents scoped, audited access, so you give your product and your agents a model call each customer pays for themselves without wiring it twice.
Used by these teams
Compare with
Often used alongside
Tools the same teams tend to run next to Hugging Face, across other categories.
Hugging Face integration FAQ
How do I add a Hugging Face integration to my product?
Enable the Hugging Face connector in your fastn dashboard, then let each customer authenticate their own Hugging Face account. fastn handles the OAuth flow, token storage and refresh per tenant, so there is no Hugging Face client code in your app and no per-customer branch in your codebase. Setup is 4 steps.
Do my customers each connect their own Hugging Face account?
Yes. Every connection is scoped to the individual customer, so each authorises their own Hugging Face account and only ever sees their own models, requests and outputs. That per-tenant isolation is the point of an embedded integration: you support the long tail of customer setups without maintaining an integration per customer.
Can AI agents use this Hugging Face integration?
Yes. The same connection is exposed to your agents through the fastn MCP gateway, with permissions scoped per tenant and every call audited. An agent runs inference on the model a customer selected and cites which endpoint answered it.
Who maintains the Hugging Face integration?
fastn does. When Hugging Face changes an endpoint, deprecates a field or alters its auth, the fix lands in the connector rather than in your backlog, and your customers' connections keep working.
Whose Hugging Face API key and quota does each call use?
Each customer authorises their own Hugging Face account, so usage, rate limits and cost land on the customer that caused them. You are not metering a shared key and re-billing it, and one heavy customer cannot exhaust another's quota.
Are inputs and outputs auditable?
Yes. Every call is logged per tenant with the call, the input and the result, so an output can be traced back to what produced it. That matters more here than in most integrations, because a generated answer or an extracted field cannot be reconstructed from the request alone.
What can I build with the Hugging Face integration?
A common starting point: let each customer bring their own Hugging Face token so inference runs on their account and against their quota. Teams also use it for the other use cases listed above, and expose it to agents for governed reads and writes.
How much does the Hugging Face integration cost?
It is included. Pricing is based on connected accounts, not on how many connectors you enable, so adding Hugging Face does not change your per-connector cost. You can start free with 3 connected accounts.
Add Hugging Face to your product
Start free with 3 connected accounts. No sales call required, and no per-customer integration code.