Skip to content

芯启神枢,NPU 先行 · Add the Intel NPU OpenVINO backend - #64

Merged
eric8810 merged 1 commit into
mainfrom
feat/openvino-npu
Sep 29, 2026
Merged

eric8810 merged 1 commit into
mainfrom
feat/openvino-npu

Conversation

@eric8810

Copy link
Copy Markdown
Contributor

Summary

Verification

Checklist

  • The change is focused; relevant tests were added or updated when behavior changed, otherwise this is N/A.
  • Documentation and CHANGELOG.md were updated for user-visible changes, otherwise this is N/A.
  • No credentials, private OCR inputs, generated build trees, or unrelated artifacts are included.

@eric8810 eric8810 changed the title feat(openvino): 芯启神枢,NPU 先行 · Add the Intel NPU OpenVINO backend 芯启神枢,NPU 先行 · Add the Intel NPU OpenVINO backend Sep 29, 2026

@eric8810 eric8810 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

审核结论

可合并。CI 全绿,无阻塞缺陷;qualification-only 的边界处理严谨,测试设计好(引擎级 openvino 测试的失败点都在 runtime 装载之前,CI 无 NPU 可跑)。以下为非阻塞意见。

值得讨论

  1. 哈希校验依赖 descriptor 自愿声明 — src/inference/openvino/backend.cpp:130:只有 descriptor 声明 bytes/sha256 才校验;只声明 openvino_runtime_library 而不带哈希时直接 dlopen,valid_runtime_policy 也允许该组合。与 WebGPU 模式一致,但文档 §6.1 写的是"装载前校验每个库的字节数与 SHA-256"。建议在 Phase B 实施状态清单补一条,或收紧为"声明 library 必须三件齐全"。
  2. 推理全程持进程级锁 — src/inference/openvino/backend.cpp:579:run() 全程持有 runtime_state().mutex,detection 与 recognition 的推理被串行化,析构(:483)也抢同一把锁。当前顺序 pipeline 无影响,但 §10 评估过的多页流水线类方案将来会撞上。锁的职责是保护 Runtime 生命周期,可考虑缩小到装载/卸载段。
  3. Auto 在无 NPU 主机也 dlopen 全套 runtime:无 /dev/accel 的机器上 Auto 仍会 dlopen OpenVINO 并创建 core,枚举设备后才以 adapter_unavailable 跳过。文档 §6.4 口径是"/dev/accel 不存在 → adapter_unavailable",实现可先做这个廉价检查再决定是否装载。
  4. run() 无 shape 合约防御:CoreML 的 run() 校验 shape 在 qualified 范围内(src/inference/coreml/backend.mm:383-390),OpenVINO 的 run() 接受任意 4 维 shape 并按需编译。detection 按实际 shape 编译是设计内的,但 recognition 传入非桶宽度(上层 bug)时会静默编译新 shape、挤占 20 桶 LRU。建议对 ModelKind::recognition 加桶校验。
  5. 桶列表双份硬编码:20 个宽度桶在 src/inference/openvino/backend.hpp:24 与 src/model/model_bundle.cpp:397(Apple bundle 校验)各写一份,将来变更需两处同步,建议共享常量或加 cross-check 断言。
  6. Node CLI 缺 npm 包错误路径测试:--provider openvino 在不含 OpenVINO 的 npm 包上应返回 unsupported_capability,目前只有 C++ 单测等价覆盖,bindings/node/test/cli.test.cjs:145 只测了非法值。建议补一条。
  7. PR 描述未填:Summary/Verification/Checklist 全是空模板;真机 14-fixture 198/199、4.8–5.4× 的数据应摘要进 Verification。

小问题

  • tests/integration/main.cpp:125:webgpu_runtime 现在涵盖 openvino,建议改名 accelerator_runtime。
  • tests/unit/test_selection.cpp:162:HAS_OPENVINO 时 early return 连 webgpu 在候选序列中间位置的断言也跳过了,可改为直接断言 openvino → webgpu → cpu 全序。
  • src/inference/openvino/backend.cpp:204:core_get_property 失败时若实现方在错误路径写入了 value 会泄漏;失败分支也可 free(value) 防御。
  • src/inference/openvino/backend.cpp:126:is_symlink 检查冗余(symlink_status + is_regular_file 已排除),无害。
  • docs/linux-device-acceleration.md:204:"CPU fallback ;"中文分号前多了空格。

已核实正确的关键点

  • 失败语义与文档一致:model_compute_unsupported 在 Auto 下 skippable、显式指定时 fatal;无 NPU → adapter_unavailable 跳过 → webgpu。
  • D-CPU 路由不隐藏 CPU:detection session 的 SessionExecutionInfo 如实报告 ORT CPU,聚合字段报 "OpenVINO" 与 Apple/MLCPU 先例一致;cpuPartition=forbid 在 D-CPU 下创建前正确拒绝。
  • natural_content_width 推导正确:content = min(clamp 后未取整宽度, ceil(48·ratio)),普通行/高瘦行/极窄/极宽四种情况都与 CPU batch-1 逐像素一致。
  • 桶列表与 Apple 锁定契约完全一致(320…3200,20 个,%32)。
  • 缓存设计合理:identity 含模型 SHA/版本/架构/驱动/compiler;锁失败静默降级为无缓存编译;prune 排除 .lock。
  • 构建边界清晰:headers-only、dlopen 装载、非 Linux x64 显式 FATAL_ERROR、npm 包行为与 CHANGELOG 一致。

@eric8810
eric8810 merged commit 0273e84 into main Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant