芯启神枢,NPU 先行 · Add the Intel NPU OpenVINO backend - #64
Merged
Merged
Conversation
eric8810
commented
Sep 29, 2026
eric8810
left a comment
Contributor
Author
There was a problem hiding this comment.
审核结论
可合并。CI 全绿,无阻塞缺陷;qualification-only 的边界处理严谨,测试设计好(引擎级 openvino 测试的失败点都在 runtime 装载之前,CI 无 NPU 可跑)。以下为非阻塞意见。
值得讨论
- 哈希校验依赖 descriptor 自愿声明 —
src/inference/openvino/backend.cpp:130:只有 descriptor 声明 bytes/sha256 才校验;只声明openvino_runtime_library而不带哈希时直接 dlopen,valid_runtime_policy也允许该组合。与 WebGPU 模式一致,但文档 §6.1 写的是"装载前校验每个库的字节数与 SHA-256"。建议在 Phase B 实施状态清单补一条,或收紧为"声明 library 必须三件齐全"。 - 推理全程持进程级锁 —
src/inference/openvino/backend.cpp:579:run()全程持有runtime_state().mutex,detection 与 recognition 的推理被串行化,析构(:483)也抢同一把锁。当前顺序 pipeline 无影响,但 §10 评估过的多页流水线类方案将来会撞上。锁的职责是保护Runtime生命周期,可考虑缩小到装载/卸载段。 - Auto 在无 NPU 主机也 dlopen 全套 runtime:无
/dev/accel的机器上 Auto 仍会 dlopen OpenVINO 并创建 core,枚举设备后才以adapter_unavailable跳过。文档 §6.4 口径是"/dev/accel不存在 → adapter_unavailable",实现可先做这个廉价检查再决定是否装载。 run()无 shape 合约防御:CoreML 的run()校验 shape 在 qualified 范围内(src/inference/coreml/backend.mm:383-390),OpenVINO 的run()接受任意 4 维 shape 并按需编译。detection 按实际 shape 编译是设计内的,但 recognition 传入非桶宽度(上层 bug)时会静默编译新 shape、挤占 20 桶 LRU。建议对ModelKind::recognition加桶校验。- 桶列表双份硬编码:20 个宽度桶在
src/inference/openvino/backend.hpp:24与src/model/model_bundle.cpp:397(Apple bundle 校验)各写一份,将来变更需两处同步,建议共享常量或加 cross-check 断言。 - Node CLI 缺 npm 包错误路径测试:
--provider openvino在不含 OpenVINO 的 npm 包上应返回unsupported_capability,目前只有 C++ 单测等价覆盖,bindings/node/test/cli.test.cjs:145只测了非法值。建议补一条。 - PR 描述未填:Summary/Verification/Checklist 全是空模板;真机 14-fixture 198/199、4.8–5.4× 的数据应摘要进 Verification。
小问题
tests/integration/main.cpp:125:webgpu_runtime现在涵盖 openvino,建议改名accelerator_runtime。tests/unit/test_selection.cpp:162:HAS_OPENVINO 时 early return 连 webgpu 在候选序列中间位置的断言也跳过了,可改为直接断言openvino → webgpu → cpu全序。src/inference/openvino/backend.cpp:204:core_get_property失败时若实现方在错误路径写入了value会泄漏;失败分支也可free(value)防御。src/inference/openvino/backend.cpp:126:is_symlink检查冗余(symlink_status+is_regular_file已排除),无害。docs/linux-device-acceleration.md:204:"CPU fallback ;"中文分号前多了空格。
已核实正确的关键点
- 失败语义与文档一致:
model_compute_unsupported在 Auto 下 skippable、显式指定时 fatal;无 NPU →adapter_unavailable跳过 → webgpu。 - D-CPU 路由不隐藏 CPU:detection session 的
SessionExecutionInfo如实报告 ORT CPU,聚合字段报 "OpenVINO" 与 Apple/MLCPU 先例一致;cpuPartition=forbid在 D-CPU 下创建前正确拒绝。 natural_content_width推导正确:content = min(clamp 后未取整宽度, ceil(48·ratio)),普通行/高瘦行/极窄/极宽四种情况都与 CPU batch-1 逐像素一致。- 桶列表与 Apple 锁定契约完全一致(320…3200,20 个,%32)。
- 缓存设计合理:identity 含模型 SHA/版本/架构/驱动/compiler;锁失败静默降级为无缓存编译;prune 排除
.lock。 - 构建边界清晰:headers-only、dlopen 装载、非 Linux x64 显式 FATAL_ERROR、npm 包行为与 CHANGELOG 一致。
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Verification
Checklist
CHANGELOG.mdwere updated for user-visible changes, otherwise this is N/A.