microsoft / microsoft/foundry-dev-tools

Foundry Local NPU Model Works in Playground but Always Results in Error 400 in the Chat Interface.

Open
#368 4 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug needs attention
Dominant language
JavaScript
Stars
2.1k
Forks
260
Avg merge
42m
Merged PRs (30d)
29

Description

Intel Ultra 7 155H
NPU Driver: 32.0.100.4621
VS Code: 1.112.0

I'm trying to utilize the local foundry NPU models (Qwen 2.5 Coder 7B and Phi 4 Mini). This is a clean install with an empty project folder, so the context should be minimal. I'm able to download the models and use them in the playground. They correctly utilize the NPU for inference. A simple Create a hello world python application. gives a good response and the following output in the AI Toolkit console.

2026-03-22 14:58:28.919 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [1400]  2026-03-22T14:58:28.9191602-03:00 Loading model:qwen2.5-coder-7b-instruct-openvino-npu:2
2026-03-22 15:00:31.213 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [1401]  2026-03-22T15:00:31.2127447-03:00 Finish loading model:qwen2.5-coder-7b-instruct-openvino-npu:2 elapsed time:00:02:02.2935276
2026-03-22 15:00:31.214 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:00:31.2133103-03:00 [Telemetry] AppName:Neutron UserAgent:NeutronServer Command:ModelLoad Status:Success Direct:True Time:122295ms
2026-03-22 15:00:31.282 [info] Agent unlocked
2026-03-22 15:00:53.135 [info] Warning: AgentRpc-pipe-#1 [0]  2026-03-22T15:00:53.1237584-03:00 EventWorker.Start error: Error Starting ETW:  Access Denied (Administrator rights required to start ETW)
2026-03-22 15:00:53.166 [info] Information: Microsoft.Neutron.OpenAI.Delegates.OpenAIApi [0]  2026-03-22T15:00:53.1662432-03:00 HandleChatCompletionAsStreamRequest -> model:qwen2.5-coder-7b-instruct-openvino-npu:2 MaxCompletionTokens:(null) maxTokens:(null) temperature:(null) topP:(null)
2026-03-22 15:00:53.168 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:00:53.1684979-03:00 Loaded cached model info for 71 models. SavedAt:3/22/2026 2:56:28 PM
2026-03-22 15:00:53.170 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:00:53.1706139-03:00 [Telemetry] AppName:Neutron UserAgent:vscode-ai-foundry/0.18.0 00cef9d4-6748-4ada-934b-1df8ad1dffb2 Command:OpenAIChatCompletions Status:Success Direct:True Time:5ms
2026-03-22 15:00:53.171 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:00:53.1709012-03:00 HandleChatCompletionAsStreamRequest -> model:qwen2.5-coder-7b-instruct-openvino-npu:2 MaxCompletionTokens:(null) maxTokens:(null) temperature:(null) topP:(null)
2026-03-22 15:00:53.171 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:00:53.1715276-03:00 AppendTokenSequences -> numTokensAppended:36
2026-03-22 15:01:01.541 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:01:01.5406485-03:00 Using new generator
2026-03-22 15:01:33.404 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:01:33.4035023-03:00 TTFT: 1 ms, TTST: 292
2026-03-22 15:01:33.405 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:01:33.4039063-03:00 Completion stats: Total Tokens/Second: 4.9, Time to First Token: 8370 ms, Average Token Generation Time: 200.8 ms, Total Time: 40232 ms, Total Tokens: 199
2026-03-22 15:01:48.529 [info] Loading View: modelLabProfilingRealtime

The same prompt in the chat interface results in:

2026-03-22 15:12:08.537 [info] Calling Foundry Local REST API: http://localhost:5272/foundry/list
2026-03-22 15:12:08.541 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:12:08.5396807-03:00 [Telemetry] AppName:Neutron UserAgent:NeutronServer Command:ListLoadedModels Status:Success Direct:True Time:0ms
2026-03-22 15:12:08.544 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:12:08.5404562-03:00 [Telemetry] AppName:Neutron UserAgent:NeutronServer Command:ListLoadedModels Status:Success Direct:True Time:0ms
2026-03-22 15:12:08.546 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:12:08.545429-03:00 Loaded cached model info for 71 models. SavedAt:3/22/2026 2:56:28 PM
2026-03-22 15:12:08.546 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:12:08.5459288-03:00 Loaded cached model info for 71 models. SavedAt:3/22/2026 2:56:28 PM
2026-03-22 15:12:08.581 [info] Information: Microsoft.Neutron.OpenAI.Provider.WindowsCopilotRuntimeServiceProvider [0]  2026-03-22T15:12:08.580958-03:00 Phi Silica model is not supported in this device
2026-03-22 15:12:08.581 [info] Information: Microsoft.Neutron.OpenAI.Provider.WindowsCopilotRuntimeServiceProvider [0]  2026-03-22T15:12:08.5811593-03:00 Phi Silica model is not supported in this device
2026-03-22 15:12:08.668 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:12:08.6666944-03:00 [Telemetry] AppName:Neutron UserAgent:NeutronServer Command:ListDownloadedModels Status:Success Direct:True Time:122ms
2026-03-22 15:12:08.668 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:12:08.6666955-03:00 [Telemetry] AppName:Neutron UserAgent:NeutronServer Command:ListDownloadedModels Status:Success Direct:True Time:121ms
2026-03-22 15:12:09.678 [info] Information: Microsoft.Neutron.AzureFoundry.AzureFoundryService [0]  2026-03-22T15:12:09.6778491-03:00 Model Phi-4-reasoning-generic-cpu:1 does not have a valid prompt template.
2026-03-22 15:12:10.082 [info] Information: Microsoft.Neutron.AzureFoundry.AzureFoundryService [0]  2026-03-22T15:12:10.0817628-03:00 Model Phi-4-reasoning-generic-gpu:1 does not have a valid prompt template.
2026-03-22 15:12:10.366 [info] Information: Microsoft.Neutron.AzureFoundry.AzureFoundryService [0]  2026-03-22T15:12:10.3652933-03:00 Total models fetched across all pages: 71
2026-03-22 15:12:10.368 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:12:10.3658188-03:00 [Telemetry] AppName:Neutron UserAgent:axios/1.13.2 Command:ModelList Status:Success Direct:True Time:1826ms
2026-03-22 15:12:10.370 [info] Information: Microsoft.Neutron.OpenAI.Provider.OpenAIServiceProviderOnnx [0]  2026-03-22T15:12:10.3693161-03:00 Loaded cached model info for 71 models. SavedAt:3/22/2026 2:56:28 PM
2026-03-22 15:12:10.401 [info] Information: Microsoft.Neutron.OpenAI.Provider.WindowsCopilotRuntimeServiceProvider [0]  2026-03-22T15:12:10.4007326-03:00 Phi Silica model is not supported in this device
2026-03-22 15:12:10.485 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:12:10.484416-03:00 [Telemetry] AppName:Neutron UserAgent:NeutronServer Command:ListDownloadedModels Status:Success Direct:True Time:115ms
2026-03-22 15:12:10.488 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:12:10.48713-03:00 [Telemetry] AppName:Neutron UserAgent:NeutronServer Command:ListLoadedModels Status:Success Direct:True Time:0ms
2026-03-22 15:12:11.071 [info] Information: Microsoft.Neutron.Telemetry.MicrosoftTelemetry [0]  2026-03-22T15:12:11.0688761-03:00 [Telemetry] AppName:Neutron UserAgent:NeutronServer Command:ModelLoad Status:Success Direct:True Time:0ms
2026-03-22 15:12:11.073 [info] Information: Microsoft.Neutron.OpenAI.Delegates.OpenAIApi [0]  2026-03-22T15:12:11.072883-03:00 HandleChatCompletionAsStreamRequest -> model:qwen2.5-coder-7b-instruct-openvino-npu:2 MaxCompletionTokens:(null) maxTokens:(null) temperature:(null) topP:(null)
2026-03-22 15:12:11.077 [error] Unable to call the qwen2.5-coder-7b-instruct-openvino-npu:2 inference endpoint due to 400.  Please check if the input or configuration is correct. 400 status code (no body) 
2026-03-22 15:12:11.078 [info] Error: Microsoft.Neutron.OpenAI.Delegates.OpenAIApi [0]  2026-03-22T15:12:11.0761834-03:00 Your input message is too large. This model supports at most 1536 completion tokens. (Parameter 'chatRequest')

This happens in both "agent" and "ask" mode. I don't see a way to view the actual input message being sent to the model or know how to set the variables MaxCompletionTokens, maxTokens, etc. when using this model.

I also receive the error

[info] [User selected model qwen2.5-coder-7b-instruct-openvino-npu:2 is not in the list of generic models: gpt-41-copilot, falling back to default model.] 

When setting this model as the option for Selected Completion Model.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the failure in the VS Code chat interface with qwen2.5-coder-7b-instruct-openvino-npu:2, then inspect the request sent through the Foundry Local REST API and the Selected Completion Model handling. Compare it with the successful playground request, focusing on completion-token limits and model fallback. Done means agent and ask modes complete without HTTP 400 and the selected model is not replaced by the default.

Written by the indexing model from the issue text.

Assessment

Tech stack
javascript, vscode
Domain
ai, devtools
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.