OrchestrAI Live

Integration · ML platform

Hugging Face + OrchestrAI

Catalog exported 2026-09-02 · Hugging Face website

Search Hugging Face models, run inference, resolve datasets, and deploy Inference Endpoints from chat.

OrchestrAI exposes 4 Hugging Face operations: 3 are low-risk (read-only or low-impact), and 1 create or modify resources and run only after you confirm the plan.

4operations
3low risk
1create or modify
0destructive
0step-level approval

What teams use it for

ML teams use OrchestrAI to search the Hugging Face Hub for a candidate model, send it a test prompt through the inference API, and deploy it to an Inference Endpoint when it performs well. Model search, inference, and dataset resolution are low risk, while deploying an endpoint is medium risk and runs on request because it creates a new billable resource. There is no operation to pause, scale, or delete an Inference Endpoint, and no model upload, so manage endpoints and repos in the Hub UI.

Every Hugging Face operation, with its risk level

Hugging Face operations available through OrchestrAI
Operation What it does Risk Step-level approval
Download HuggingFace Dataset Resolve a HuggingFace dataset for download Low risk No
HuggingFace Inference Run inference against a HuggingFace model Low risk No
List HuggingFace Models Search/list models on the HuggingFace Hub Low risk No
Deploy HuggingFace Model Deploy a model to a HuggingFace Inference Endpoint Creates resources No

Risk tiers come from the catalog: low is read-only or low-impact, medium creates resources and is reversible, high modifies existing resources, destructive may lose data. Every plan that creates or changes resources is shown with its cost estimate and waits for your confirmation. Operations marked with a step-level approval pause again on their own step. Destructive operations require a typed risk phrase.

What you connect

A Hugging Face credential (stored as huggingface). Connected-service tokens are envelope-encrypted with a per-record key wrapped by a cloud KMS.

Prompts that work

  • Search Hugging Face for text-classification models under 500M parameters sorted by downloads
  • Run inference on sentence-transformers/all-MiniLM-L6-v2 with the sentence refund my order
  • Deploy meta-llama/Llama-3.1-8B-Instruct to an Inference Endpoint on an A10G in us-east-1

Before anything runs

Every mutation shows its plan, cost estimate, and blast radius, then waits for your confirmation. Destructive operations require a typed risk phrase. Credentials are minted per run through OIDC federation and discarded afterward; nothing you create here is invisible later, because every resource lands in the desired-state ledger where drift is detected and can be converged. Details on the security page.

Frequently asked questions

Can OrchestrAI delete a Hugging Face Inference Endpoint?
No. Deploying an endpoint is available as a medium-risk operation, but pausing, scaling, and deleting endpoints are not, so use the Hugging Face UI for those.
Does running Hugging Face inference through OrchestrAI require approval?
No. Inference is low risk and changes nothing in your infrastructure, so it runs immediately with your Hugging Face token.
How does OrchestrAI authenticate to Hugging Face?
You add a Hugging Face credential once in the connections screen. It is envelope-encrypted with a per-record key wrapped by a cloud KMS and is only decrypted inside the run that needs it.

Related integrations

Try it on your own account

Connect your cloud read-only and see your resources, drift, and costs before anything runs. $5 minimum to start. Unused credits refunded in your first 14 days.

Start for $5

Unused credits refunded in your first 14 days.