forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 10
Pull requests: ravi9/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Optimize NPU performance with KV cache slicing and requantization
#318
opened Sep 15, 2026 by
zhaixuejun1993
Collaborator
Loading…
Enhance NPU optimizations with KV cache slicing and buffer management
#314
opened Sep 8, 2026 by
zhaixuejun1993
Collaborator
Loading…
ggml-openvino : make op support a property of the registry
#307
opened Sep 2, 2026 by
cavusmustafa
Collaborator
•
Draft
ggml-openvino : fix 2D/3D view input shape inference
#303
opened Aug 27, 2026 by
mostafafaheem
Collaborator
•
Draft
[Draft] Refactor OpenVINO backend for better modularity and documentation
#279
opened Aug 11, 2026 by
zhaixuejun1993
Collaborator
Loading…
Refactor OpenVINO backend: isolate op support, manage buffers, and document API
#278
opened Aug 10, 2026 by
zhaixuejun1993
Collaborator
Loading…
OpenVINO: PRD-compliant device enumeration and memory reporting for --list-devices
#256
opened Jul 16, 2026 by
haarika-madaka
Loading…
Refactor OpenVINO backend: move and clean up parameter handling
#240
opened Jul 3, 2026 by
zhaixuejun1993
Collaborator
Loading…
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.