> ## Documentation Index
> Fetch the complete documentation index at: https://geniex.aihub.qualcomm.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Local server

> Run an OpenAI-compatible HTTP API on localhost backed by Snapdragon NPU/GPU/CPU acceleration.

GenieX includes a built-in inference server that exposes an **OpenAI-compatible API**. Run models on-device and connect them to any application or framework that speaks the OpenAI protocol — agentic frameworks like **LangChain**, AI-native apps like **OpenClaw**, or your own code. No cloud dependency.

## **Prerequisites**

* The CLI installed — see [Install](/en/run/cli/install).
* Interactive shell from container (Docker only) — see [Run interactively](/en/run/linux/install#run-interactively).
* A model pulled. `geniex serve` does **not** auto-download models.

## **Start the server**

Pull a model:

```bash bash theme={"dark"}
geniex pull ai-hub-models/Qwen3-4B-Instruct-2507
```

Start the server:

```bash bash theme={"dark"}
geniex serve
```

The server runs on `http://127.0.0.1:18181` by default. Keep this terminal open and make requests from another one. Run `geniex serve -h` for all configurable options.

## **POST /v1/chat/completions**

Creates a model response for a conversation. Supports LLM (text-only) and VLM (text + image or audio).

### **LLM request**

```json Example Value theme={"dark"}
{
  "model": "ai-hub-models/Qwen3-4B-Instruct-2507",
  "messages": [
    {"role": "user", "content": "Hello! Briefly introduce yourself."}
  ],
  "max_tokens": 256,
  "temperature": 0.7,
  "stream": false
}
```

### **Try it from Swagger UI**

Open `http://127.0.0.1:18181` in your browser to access the built-in Swagger UI.

**Step 1.** Expand the `POST /v1/chat/completions` endpoint to view the example request body and schema.

<img src="https://mintcdn.com/qualcomm-0801e48b/VijZ6eXFSGIGNc9Z/Images/server/curl-1.png?fit=max&auto=format&n=VijZ6eXFSGIGNc9Z&q=85&s=02340eb8c5b27abce2cfd1bc5a5541b0" alt="Swagger UI showing the chat completions endpoint with example request body" width="1376" height="1059" data-path="Images/server/curl-1.png" />

**Step 2.** Click **Try it out**, edit the request body as needed, then click **Execute**.

<img src="https://mintcdn.com/qualcomm-0801e48b/VijZ6eXFSGIGNc9Z/Images/server/curl-2.png?fit=max&auto=format&n=VijZ6eXFSGIGNc9Z&q=85&s=588ba0820785188699e9b2885cbc7ece" alt="Editing the request body in Try-it-out mode before executing" width="1387" height="1447" data-path="Images/server/curl-2.png" />

**Step 3.** View the response — a `200` status with the model's generated reply.

<img src="https://mintcdn.com/qualcomm-0801e48b/VijZ6eXFSGIGNc9Z/Images/server/curl-3.png?fit=max&auto=format&n=VijZ6eXFSGIGNc9Z&q=85&s=e76944247469636be79fe528bee9ec1e" alt="Response body showing a successful 200 response with the model's reply" width="1353" height="737" data-path="Images/server/curl-3.png" />

### **VLM request**

A VLM accepts multimodal content parts alongside text. Use `image_url` for images, and `input_audio` for audio on a model whose mmproj carries a conformer encoder (e.g. `google/gemma-4-E2B-it-qat-q4_0-gguf`). Both `image_url.url` and `input_audio.data` accept the same three formats:

| Format                                             | Image example                          | Audio example                          |
| -------------------------------------------------- | -------------------------------------- | -------------------------------------- |
| Local file path (the `file://` prefix is optional) | `C:/Users/Username/Pictures/photo.jpg` | `/data/jfk.wav`, `file:///tmp/jfk.wav` |
| HTTP / HTTPS URL — fetched by the server           | `https://example.com/image.jpg`        | `https://example.com/clip.mp3`         |
| Base64 data URL — inline bytes                     | `data:image/png;base64,iVBORw0KGgo...` | `data:audio/wav;base64,UklGR...`       |

<Note>
  **Running in Docker?** Local paths are resolved **inside the container**, not on your host. The install command already mounts `$PWD/data` to `/data` — drop your images and audio files there and pass `/data/cat.jpg` / `/data/jfk.wav`. Alternatively, use an HTTP URL or base64 data URL to skip the filesystem entirely.
</Note>

<Warning>Audio input runs on the **llama.cpp** backend only. QAIRT models report `audio: false` and a QAIRT model given audio fails with `GenieXError(-201201): Multimodal generation failed`.</Warning>

A single message can mix `image_url` and `input_audio` parts. Pull an audio-capable model and grab a sample clip:

```bash theme={"dark"}
geniex pull google/gemma-4-E2B-it-qat-q4_0-gguf
curl -L -o jfk.wav https://github.com/ggml-org/whisper.cpp/raw/master/samples/jfk.wav
curl -L -o landmark.jpg "https://images.pexels.com/photos/402028/pexels-photo-402028.jpeg?w=1024"
```

```bash theme={"dark"}
curl http://127.0.0.1:18181/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemma-4-E2B-it-qat-q4_0-gguf",
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "text", "text": "Describe the image, then transcribe the audio."},
          {"type": "image_url", "image_url": {"url": "/full/path/to/landmark.jpg"}},
          {"type": "input_audio", "input_audio": {"data": "/full/path/to/jfk.wav"}}
        ]
      }
    ],
    "max_tokens": 256
  }'
```

The same request from the Python `openai` client:

```python theme={"dark"}
resp = client.chat.completions.create(
    model="google/gemma-4-E2B-it-qat-q4_0-gguf",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Describe the image, then transcribe the audio."},
                {"type": "image_url", "image_url": {"url": "/full/path/to/landmark.jpg"}},
                {"type": "input_audio", "input_audio": {"data": "/full/path/to/jfk.wav"}},
            ],
        }
    ],
    max_tokens=256,
)
print(resp.choices[0].message.content)
```

Output (verified on `--compute npu`, Snapdragon X Elite):

```text theme={"dark"}
This is a photograph featuring a traditional Japanese temple ... In the background, there are softer, blue-toned mountains and a distant urban skyline.

**Transcription:**
And so my fellow Americans ask not what your country can do for you, ask what you can do for your country.
```

In Swagger UI, replace the request body with a VLM payload like the one above, point the paths at local files, then click **Execute**.

<img src="https://mintcdn.com/qualcomm-0801e48b/VijZ6eXFSGIGNc9Z/Images/server/curl-4.png?fit=max&auto=format&n=VijZ6eXFSGIGNc9Z&q=85&s=c41c6a3a155739709d96ff355e832ee9" alt="Editing the VLM request body in Try-it-out mode before executing" width="1871" height="968" data-path="Images/server/curl-4.png" />

<img src="https://mintcdn.com/qualcomm-0801e48b/VijZ6eXFSGIGNc9Z/Images/server/curl-5.png?fit=max&auto=format&n=VijZ6eXFSGIGNc9Z&q=85&s=e01e48ddc6441fd8bbe23bb656482e22" alt="Response body showing a successful 200 response with the VLM's image description" width="1844" height="634" data-path="Images/server/curl-5.png" />

## **POST /v1/completions**

Generates a raw continuation of `prompt` with no chat template applied — the prompt reaches the model verbatim. Use this for clients that build the full prompt themselves, such as editor fill-in-the-middle (FIM) code autocompletion: put the model's FIM tokens directly in `prompt`. LLM models only.

```json Example Value theme={"dark"}
{
  "model": "unsloth/Qwen2.5-Coder-3B-GGUF:Q4_0",
  "prompt": "<|fim_prefix|>def fibonacci(n):\n    <|fim_suffix|>\n    return a<|fim_middle|>",
  "max_tokens": 64,
  "temperature": 0.2,
  "stop": ["<|endoftext|>"],
  "stream": false
}
```

The response `choices[0].text` is the raw completion, ready to insert at the cursor. `stream`, `stop`, `echo` and the sampler knobs (`temperature`, `top_p`, `top_k`, `min_p`, `repetition_penalty`, `seed`) work as on `/v1/chat/completions`; `suffix` is not supported — encode the suffix with the model's FIM tokens inside `prompt` instead.

## **Python client (OpenAI SDK)**

Because the server speaks the OpenAI protocol, you can point the official `openai` Python client at the local endpoint and reuse any existing OpenAI code. Install with `pip install openai`, then create a client:

```python python theme={"dark"}
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:18181/v1",
    api_key="geniex",  # any non-empty string; the server does not check it
)
```

The examples below reuse this `client`. Replace the `model` value with a model you have already pulled. The optional `:<precision>` suffix (e.g. `Q4_0`, `Q4_K_M`, `Q8_0`) selects a quantization variant — `Q4_0` is recommended for llama.cpp on Hexagon NPU. See [Precisions (Quantizations) Supported](/en/models/supported#precisions-quantizations-supported).

### **Streaming**

Print each delta as it arrives:

```python python theme={"dark"}
stream = client.chat.completions.create(
    model="unsloth/Qwen3-4B-GGUF:Q4_0",
    messages=[
        {"role": "user", "content": "Hello! Briefly introduce yourself."},
    ],
    max_tokens=256,
    temperature=0.7,
    stream=True,
)

for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)
print()
```

### **Chat completion (non-streaming)**

Single request, single response, no streaming — the standard OpenAI `chat.completions.create` shape. The `enable_think=False` extra parameter turns off Qwen3's default `<think>…</think>` reasoning prefix so the reply content stays clean.

```python python theme={"dark"}
resp = client.chat.completions.create(
    model="unsloth/Qwen3-4B-GGUF:Q4_0",
    messages=[
        {"role": "user", "content": "Hello! Briefly introduce yourself."},
    ],
    max_tokens=128,
    temperature=0.7,
    extra_body={"enable_think": False},
)

print(resp.choices[0].message.content)
print("finish_reason:", resp.choices[0].finish_reason)
print("usage:", resp.usage)
```

Output:

```text theme={"dark"}
Hello! I'm Qwen, a large language model developed by Alibaba Cloud. I can help with a wide range of tasks, including answering questions, writing articles, creating stories, and more. I'm here to assist you in any way I can! How can I help you today?
finish_reason: stop
usage: CompletionUsage(completion_tokens=58, prompt_tokens=19, total_tokens=77, ...)
```

### **Separating reasoning (reasoning\_content)**

By default a thinking model leaves its chain-of-thought inline in `message.content` (`<think>…</think>`), mixed in with the final reply. Pass `reasoning_format="deepseek"` to make the server move the chain-of-thought into the OpenAI-standard `message.reasoning_content` field, leaving `content` clean.

<Note>
  `reasoning_format` is not the same as `enable_think`: `enable_think=False` makes the model **not produce** a chain-of-thought at all, while `reasoning_format="deepseek"` lets it think as usual but **moves** the thinking out of `content`. Values are `none` (default, kept inline) and `deepseek` / `deepseek-legacy` / `auto` (all separate). Tool-call requests ignore this parameter (tool parsing needs the raw tagged text).
</Note>

```python python theme={"dark"}
resp = client.chat.completions.create(
    model="unsloth/Qwen3-4B-GGUF:Q4_0",
    messages=[
        {"role": "user", "content": "What is 2+2? Answer briefly."},
    ],
    max_tokens=128,
    extra_body={"reasoning_format": "deepseek"},
)

msg = resp.choices[0].message
print("reasoning:", msg.model_extra.get("reasoning_content"))
print("content:", msg.content)
```

Streaming works the same way — the chain-of-thought arrives as `delta.reasoning_content` deltas and the final reply as `delta.content`.

### **Tool calling**

Function/tool calling uses the standard OpenAI `tools` schema. The server extracts the tool call from the model's generated text (`<tool_call>…</tool_call>` tags or a fenced ` ```json ` block) and re-emits it as OpenAI `tool_calls`. The flow works with **VLMs too** — the model can look at an image, decide what to search for, and call a tool.

The example below walks through a two-step agentic loop with `qualcomm/Qwen3-VL-4B-Instruct`: (1) the VLM identifies a landmark from a photo and calls `web_search`, (2) you execute the search locally and feed the results back so the VLM writes a grounded reply.

<Note>
  Only one tool call per assistant turn is parsed — parallel tool calls in a single response are not supported.
</Note>

<Note>
  Two Qwen3-VL specifics for reliable tool calls:

  1. Prime the model with a system message that spells out the `<tool_call>…</tool_call>` shape (Qwen3-VL's chat template does not enforce it as strongly as Qwen3's text-only template).
  2. On the follow-up turn, drop `tools=` and drop the image content from `messages` — this stops the VLM from re-invoking the tool and avoids re-running the vision encoder on the same image.
</Note>

Install the search library used by the tool (`pip install ddgs` — DuckDuckGo, no API key required), then:

```python python theme={"dark"}
import json

from ddgs import DDGS

tools = [
    {
        "type": "function",
        "function": {
            "name": "web_search",
            "description": "Search the web for travel information about a location.",
            "parameters": {
                "type": "object",
                "properties": {
                    "query": {"type": "string", "description": "Search query, e.g. 'things to do in Kyoto'"},
                },
                "required": ["query"],
            },
        },
    }
]

def web_search(query: str) -> list[dict]:
    return [
        {"title": r["title"], "snippet": r["body"], "url": r["href"]}
        for r in DDGS().text(query, max_results=3)
    ]

IMAGE_URL = "https://images.pexels.com/photos/402028/pexels-photo-402028.jpeg?w=1024"

system_prompt = (
    "You are a travel assistant. Call the web_search tool to look up any location "
    "the user asks about before answering. Once you receive the tool results, do not "
    "call the tool again - use them to write a short, friendly reply for the user. "
    "Emit tool calls in the exact format: "
    '<tool_call>{"name": "web_search", "arguments": {"query": "<your query>"}}</tool_call>'
)

messages = [
    {"role": "system", "content": system_prompt},
    {
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": IMAGE_URL}},
            {"type": "text", "text": "Identify the landmark, then call web_search for travel tips about it."},
        ],
    },
]

# Step 1 - VLM identifies the landmark and requests a web_search call.
first = client.chat.completions.create(
    model="qualcomm/Qwen3-VL-4B-Instruct",
    messages=messages,
    tools=tools,
    tool_choice="auto",
    max_tokens=512,
    extra_body={"enable_think": False},
)
call = first.choices[0].message.tool_calls[0]
print("finish_reason:", first.choices[0].finish_reason)  # -> "tool_calls"
print("call:", call.function.name, call.function.arguments)

# Step 2 - run the tool, feed the result back as a fresh text-only conversation.
result = web_search(**json.loads(call.function.arguments))

followup = [
    {"role": "system", "content": system_prompt},
    {"role": "user", "content": "Summarize travel tips using the search results below."},
    first.choices[0].message,
    {"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)},
]

final = client.chat.completions.create(
    model="qualcomm/Qwen3-VL-4B-Instruct",
    messages=followup,
    max_tokens=256,
    extra_body={"enable_think": False},
)
print(final.choices[0].message.content)
```

Output (grounded on the live DuckDuckGo results — the exact prose varies with what the web returns that day):

```text theme={"dark"}
finish_reason: tool_calls
call: web_search {"query": "Kiyomizu-dera travel tips"}
Here's a quick summary of travel tips for Kiyomizu-dera in Kyoto:

**1. Best Times to Visit:**
Early morning (around 7-8 AM) or late afternoon (4-6 PM) are ideal to avoid the crowds.
Peak hours (11 AM–2 PM) can be very busy, so plan accordingly.

**2. How to Get There:**
Take the Kiyomizu-dera train or bus from Kyoto Station to the temple.
The temple is located on a hill with a wooden stage built without nails, and the walkways are accessible, though narrow.

**3. What to See:**
- The great wooden stage (without nails)
- Otawa waterfall and its three streams
- Jishu love shrine
- The Sannenzaka and Ninenzaka approach streets
- Night illuminations (best experienced at night)
...
```

## **Other endpoints**

* `GET /v1/models` — list available models.
* `GET /v1/models/{model}` — get info about a specific model.

<br />

<div class="feedback-wrapper">
  <span class="feedback-label">Was this page helpful?</span>

  <div class="feedback-toggle">
    <input type="radio" name="feedback" id="feedback-yes" class="feedback-input" />

    <label for="feedback-yes" class="feedback-button">
      <img src="https://mintlify.s3.us-west-1.amazonaws.com/qualcomm-0801e48b/Images/FeedBack/thumbs-up.svg" alt="Thumbs up" class="feedback-icon" noZoom />

      Yes
    </label>

    <input type="radio" name="feedback" id="feedback-no" class="feedback-input" />

    <label for="feedback-no" class="feedback-button">
      <img src="https://mintlify.s3.us-west-1.amazonaws.com/qualcomm-0801e48b/Images/FeedBack/thumbs-down.svg" alt="Thumbs down" class="feedback-icon" noZoom />

      No
    </label>
  </div>
</div>
