Local LLM
Local LLM is Microsoft's on-device large language model system for Windows Copilot+ devices that processes AI tasks privately without cloud dependency.
Why It Exists
Users need AI assistance for writing and productivity tasks without uploading their data to the cloud for privacy and security reasons.
How It Works
Local LLM is Microsoft's on-device large language model system for Windows Copilot+ PCs. Using the device's NPU, the system runs compact language models locally to perform tasks like text summarization, rewriting, and question answering. This enables AI assistance without sending data to the cloud, ensuring maximum privacy. Local LLM supports basic productivity tasks and can escalate to cloud Copilot for complex queries when the user chooses.
Everyday Use Cases
- Text summarization
- Writing assistance
- Question answering
- Content rewriting
- Idea generation
- Offline productivity
- Private AI assistance
User Workflow
- User activates Copilot with local processing mode
- On-device LLM is loaded into NPU memory
- User enters query or text
- AI processes locally and generates response
- User can escalate to cloud Copilot if needed
AI Processing Flow
User input and context are processed entirely on-device by NPU-accelerated LLM models. No data leaves the device for basic tasks. The system uses distilled models optimized for the NPU. Model weights are stored securely on-device. For enhanced capabilities, the system can optionally connect to Microsoft's cloud Copilot with explicit user consent.
Inputs / Outputs
Inputs:
- User query
- Text content
- App context
- Document data
- User preferences
- System information
Outputs:
- AI-generated responses
- Summarized text
- Rewritten content
- Answered questions
- Generated ideas
- Private results
Known Limitations
- On-device models are less capable than cloud LLMs
- Limited context window
- Model loading takes 2-5 seconds
- Requires NPU with 16+ GB VRAM
- Some languages not supported
Unsupported Scenarios
- Multi-document analysis
- Code generation
- Data visualization
- Real-time collaboration
- Custom model loading
Performance Notes
On-device model loads in 3-5 seconds; Queries respond in 1-3 seconds; NPU utilization 80-90 percent during active processing; No internet required; Privacy guaranteed
Available On
Shows where this feature is available and how its AI processing works on each platform. Availability may vary by device.
| Platform | Execution | Offline | Cloud | OS / Software |
|---|---|---|---|---|
| Copilot+ | Local | Yes | No - on-device | Windows 11 24H2 |
Research Status
Confidence and verification reflect how complete documentation is. Fields may show Not assessed or Not yet verified while research is ongoing - this flags gaps, not product deficiencies.
| Research Status | Verified |
| Confidence | High |
| First introduced | 2024-09-04 |
| Last updated | 2026-08-16 |
| Last verified | 2026-08-16 |
Research Notes
Available on Copilot+ PCs with Snapdragon X Elite or Intel Lunar Lake; Requires 16GB RAM and NPU; Windows 11 2024 Update+; Introduced 2024
Details
| Vendor | Microsoft |
| AI Platform | Copilot+ |
| Category | Writing |
Capability Mapping
This feature maps to canonical capability family: Text Summarization, Writing.
Canonical capabilities: Text Rewrite, Text Summarization.
- Text Rewrite - Restates or rewrites existing text in a new style or tone, including AI-assisted response suggestions.
- Text Summarization - Condenses text, messages, documents, or conversations into a shorter summary.
Grounded in 1 device(s) on Copilot+.
Related Features
- AI Assistant - Complementary
This page describes platform-level capability; not device-level support for any specific product.