> ## Documentation Index
> Fetch the complete documentation index at: https://docs.beam.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Hosted Model Inference

> Discover and call models from the Endpoints catalog

Use hosted models from the dashboard's **Endpoints** catalog.

## Discover Models

```bash theme={null}
curl --fail-with-body https://app.beam.cloud/v1/models \
  -H "Authorization: Bearer $BEAM_TOKEN"
```

The catalog is public. Include a token to see models available to your workspace.

## Call a Chat Model

Install the OpenAI Python client and set `BEAM_MODEL` to a chat model ID from the catalog:

```python theme={null}
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["BEAM_TOKEN"],
    base_url="https://app.beam.cloud/v1",
)
response = client.chat.completions.create(
    model=os.environ["BEAM_MODEL"],
    messages=[{"role": "user", "content": "Explain what a container is in one sentence."}],
)
print(response.choices[0].message.content)
```

Pass `stream=True` for streaming responses. Supported options vary by model.

## Other Model Types

Use `/v1/embeddings` for embeddings or `/v1/models/{model-id}/invoke` for models with custom schemas. The model's catalog page includes request examples, image routes, and pricing.

Inference requires a workspace token and credits for prepaid workspaces. Keep the token on your server.

To deploy your own model, use an [endpoint](/v2/endpoint/overview) or [container](/v2/hosting/containers).
