Skip to content
This repository was archived by the owner on Jun 12, 2026. It is now read-only.
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
35 commits
Select commit Hold shift + click to select a range
6c31072
refactor: drop phi-3.5
swernerx Jan 12, 2026
5fd930a
refactor: drop llama-3.2 and phi-3*
swernerx Jan 12, 2026
f9e6a08
feat(hf2swift): add full GPT-OSS MoE generator support
swernerx Jan 12, 2026
4ad9d29
refactor(swift): port core infrastructure from mlx-lm Python
swernerx Jan 12, 2026
8602158
test(swift): add SwitchLayers tests
swernerx Jan 12, 2026
0fc1050
refactor(swift): clean cut - delete all infrastructure for fresh port
swernerx Jan 12, 2026
6ff9236
feat(swift): port core infrastructure from mlx-lm Python
swernerx Jan 12, 2026
beb8074
fix(generator): update hf2swift for new ported API conventions
swernerx Jan 12, 2026
8457a93
test(swift): add unit tests for ported infrastructure
swernerx Jan 12, 2026
f7117b7
docs: add mlx-lm git hash to all ported files
swernerx Jan 12, 2026
4d0971b
fix(tests): fix all unit tests to pass on Apple Silicon
swernerx Jan 12, 2026
1458f48
ci: enable Swift tests with proper Metal library setup
swernerx Jan 12, 2026
311a830
style: format documentation files
swernerx Jan 12, 2026
dee9876
fix: remove unnecessary optional chain
swernerx Jan 12, 2026
b018cef
chore: regenerate all models and fix pre-push hook path
swernerx Jan 12, 2026
4b7907e
fix: ensure consistent swiftformat in pre-push hook
swernerx Jan 12, 2026
1cc0001
docs: add PR description template and slash command
swernerx Jan 12, 2026
7795578
chore: remove temporary PR description file
swernerx Jan 12, 2026
5d4e86f
refactor(swift): extract shared model components
swernerx Jan 12, 2026
1ce56d1
refactor(swift): port GemmaRMSNorm from Python to ported/
swernerx Jan 12, 2026
5020d2f
refactor(hf2swift): separate architectural features from config values
swernerx Jan 12, 2026
1676b18
refactor: extract FusedQKVAttention into shared component
swernerx Jan 12, 2026
59ae353
refactor: extract MoESanitizer and MathUtils into shared
swernerx Jan 12, 2026
8babf8c
refactor: use shared components for simple models (Llama, Qwen2)
swernerx Jan 12, 2026
a391e6b
refactor: extract AltUp, Laurel and math utilities into shared
swernerx Jan 12, 2026
72aaa5f
refactor: extract model definitions into separate files
swernerx Jan 12, 2026
e71b1a9
docs: comprehensive documentation update for new architecture
swernerx Jan 12, 2026
dd7f945
fix: update ignore list
swernerx Jan 12, 2026
2e5b8e0
test: add comprehensive tests ported from mlx-lm
swernerx Jan 12, 2026
5d08577
chore: add .venv to gitignore and remove local venv
swernerx Jan 12, 2026
86108f3
fix(ci): ignore generated fumadocs .source/ in prettier
swernerx Jan 12, 2026
52bde4a
chore: remove old rfs
swernerx Jan 12, 2026
4ce620f
fix(ci): resolve Swift build warnings and cache issues
swernerx Jan 12, 2026
bc14aac
test: restore tests from main branch
swernerx Jan 12, 2026
468f671
fix(ci): invalidate swift cache and ensure test bundle directory exists
swernerx Jan 12, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
79 changes: 79 additions & 0 deletions .cursor/prompts/create-pr-description.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
# Create Pull Request Description

Generate a concise, benefit-focused PR description in US English.

## Guidelines

### Structure

```markdown
## Summary

[One paragraph explaining WHAT changed and WHY it matters]

## Key Changes

- [Bullet points of significant changes - focus on impact, not implementation details]

## Architecture Decisions

[Only include if there are decisions other contributors should be aware of]

## Breaking Changes

[Only include if there are breaking changes]
```

### Writing Style

- **Be concise**: Every sentence should add value
- **Focus on benefits**: What does this enable? What problem does it solve?
- **Avoid redundancy**: Don't repeat information, don't state the obvious
- **Skip boilerplate**: No "This PR adds...", no test mentions (CI handles that)
- **Use active voice**: "Ports X from Y" not "X was ported from Y"

### What to Include

- Significant architectural changes
- New capabilities or features
- Performance improvements with context
- Migration guidance if needed
- Links to related issues/RFCs

### What to Exclude

- Test coverage details (CI shows this)
- Obvious file changes (reviewers can see the diff)
- Implementation minutiae
- Changelog-style lists of every file touched

## Example

```markdown
## Summary

Switches MLX infrastructure from vendored mlx-swift-lm to direct ports from mlx-lm (Python). This gives us access to the latest model architectures faster, as mlx-lm releases more frequently and has broader model coverage.

## Key Changes

- Direct Python→Swift ports for KVCache, RoPE, and MoE layers
- New `ported/` directory structure with version tracking
- Generator now produces code matching mlx-swift-lm patterns exactly

## Architecture Decisions

**Why port from Python instead of using mlx-swift-lm?**
mlx-lm (Python) is the primary source, updated more frequently, and supports models like Llama 4 MoE that mlx-swift-lm doesn't yet have.

**Directory structure**:

- `generated/models/` - hf2swift generator output
- `ported/` - LLM-assisted ports from Python with git hash tracking
```

## Instructions

1. Analyze the current branch changes using `git log` and `git diff`
2. Read any relevant RFCs or decision documents
3. Generate a PR description following the structure above
4. Keep total length under 500 words
243 changes: 243 additions & 0 deletions .cursor/prompts/port-python-to-swift.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,243 @@
# Port Python mlx-lm to Swift

You are porting Python code from Apple's `mlx-lm` library to Swift for the `node-mlx` project.

## Source Repository

- **Primary**: https://github.com/ml-explore/mlx-lm/tree/main/mlx_lm/models
- **Reference only**: https://github.com/ml-explore/mlx-swift-lm

**IMPORTANT**: Always record the exact git hash. Get it with:

```bash
curl -s "https://api.github.com/repos/ml-explore/mlx-lm/commits/main" | grep '"sha"' | head -1
```

## File Locations

| Type | Directory |
| ----------------- | ------------------------------------------------------ |
| Ported code | `packages/swift/Sources/NodeMLXCore/ported/` |
| Shared components | `packages/swift/Sources/NodeMLXCore/shared/` |
| Tests | `packages/swift/Tests/NodeMLXCoreTests/` |
| Generated models | `packages/swift/Sources/NodeMLXCore/generated/models/` |

## Core Principles

### 1. Clean Cut Philosophy

- Start fresh, don't patch existing code
- Port with understanding, not blind translation
- Premium architect-level Swift: idiomatic, elegant, maintainable

### 2. Focus on Popular Models

| Priority | Models | Notes |
| ------------ | ----------------------------------------- | ------------------------- |
| ✅ Essential | Llama, Qwen, Phi, Gemma, Mistral, GPT-OSS | Mainstream |
| ⏸️ Defer | Mamba, Jamba, DBRX | SSM/unusual architectures |
| ❌ Skip | Batch processing, server features | Not needed for inference |

### 3. Minimal Viable Port

- Port core functionality, not edge cases
- Skip features that < 5% of users need
- Add extensibility points for future additions

## File Header Template

Every ported file **must** include:

```swift
// Copyright © 2024 Sebastian Software GmbH. All rights reserved.
// SPDX-License-Identifier: MIT
//
// Ported from mlx-lm (https://github.com/ml-explore/mlx-lm)
// Original: mlx_lm/models/<filename>.py
// Git Hash: <full-40-char-hash> (<YYYY-MM-DD>)
```

## Swift Style Guide

### Naming Conventions

| Python | Swift |
| ---------------------- | --------------------- |
| `snake_case` | `camelCase` |
| `class KVCache` | `class KVCache` |
| `def update_and_fetch` | `func updateAndFetch` |
| `__init__` | `init` |
| `__len__` | `var count: Int` |
| `_private_method` | `private func method` |

### Type Mappings

| Python | Swift |
| ------------- | ----------------- |
| `mx.array` | `MLXArray` |
| `nn.Module` | `Module` (MLXNN) |
| `Optional[T]` | `T?` |
| `List[T]` | `[T]` |
| `Dict[K, V]` | `[K: V]` |
| `Tuple[A, B]` | `(A, B)` |
| `None` | `nil` |
| `@property` | computed property |

### MLX Operations

| Python | Swift |
| -------------------------------- | ------------------------------- |
| `mx.zeros(shape)` | `MLXArray.zeros(shape)` |
| `mx.concatenate([a, b], axis=2)` | `concatenated([a, b], axis: 2)` |
| `mx.quantize(x, ...)` | `MLX.quantized(x, ...)` |
| `x[..., :n, :]` | `x[.ellipsis, ..<n, 0...]` |
| `x.shape[0]` | `x.dim(0)` |
| `x.dtype` | `x.dtype` |

## Code Structure

### Protocol-First Design

```swift
/// Protocol for all KV cache implementations
public protocol KVCacheProtocol: AnyObject {
func update(keys: MLXArray, values: MLXArray) -> (MLXArray, MLXArray)
var offset: Int { get }
func makeMask(queryLength: Int, windowSize: Int?) -> MLXFast.ScaledDotProductAttentionMaskMode
}
```

### Class Structure

```swift
/// KV cache with grow-in-place strategy
public class StandardKVCache: KVCacheProtocol {
// MARK: - Properties

private var keys: MLXArray?
private var values: MLXArray?
public private(set) var offset: Int = 0

public static let step = 256

// MARK: - Initialization

public init() {}

// MARK: - Cache Operations

public func update(keys: MLXArray, values: MLXArray) -> (MLXArray, MLXArray) {
// Implementation
}
}
```

## What NOT to Port

### From cache.py

- ❌ `BatchKVCache`, `BatchRotatingKVCache` - Server/batch processing
- ❌ `MambaCache`, `ArraysCache` - SSM models
- ❌ `ChunkedKVCache`, `CacheList` - Specialized use cases
- ❌ `save_prompt_cache`, `load_prompt_cache` - Serialization

### General

- ❌ Batch processing features
- ❌ Prompt caching to disk
- ❌ Speculative decoding caches
- ❌ Multi-modal (initially)

## Shared Components

Before porting, check if a shared component already exists in `shared/`:

| Component | File | Use When |
| ----------------- | --------------------------- | -------------------------- |
| RMSNorm | `RMSNorm.swift` | Standard RMS normalization |
| GemmaRMSNorm | `ported/GemmaRMSNorm.swift` | (1+weight) scaling |
| StandardAttention | `StandardAttention.swift` | Basic GQA attention |
| StandardMLP | `StandardMLP.swift` | SwiGLU MLP |
| MathUtils | `MathUtils.swift` | erfinv, clipResidual, topK |

## Testing

### Test File Location

Tests go in `packages/swift/Tests/NodeMLXCoreTests/`:

```swift
import XCTest
@testable import NodeMLXCore
import MLX

final class KVCacheTests: XCTestCase {
func testUpdateAndFetch() {
let cache = StandardKVCache()
let keys = MLXArray.zeros([1, 4, 8, 64])
let values = MLXArray.zeros([1, 4, 8, 64])

let (k, v) = cache.update(keys: keys, values: values)

XCTAssertEqual(cache.offset, 8)
XCTAssertEqual(k.dim(2), 8)
}
}
```

### Running Tests

```bash
cd packages/swift
swift test
```

## Workflow

1. **Download Python source**:

```bash
curl -s "https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/models/<file>.py" -o /tmp/<file>.py
```

2. **Analyze**: Essential vs. optional features

3. **Check shared components**: Reuse if exists

4. **Design Swift API**: Protocols, classes

5. **Implement**: Premium Swift patterns

6. **Test**: Comprehensive coverage

7. **Document**: Update PORTING_DECISIONS.md

8. **Build**:
```bash
cd packages/swift && swift build -c release && swift test
```

## Documentation Updates

After porting, update:

1. **File header**: Git hash, date
2. **PORTING_DECISIONS.md**: What was ported, decisions made
3. **ported/README.md**: Add to ported files table
4. **Tests**: Add test file

## Quick Reference

```bash
# Get latest mlx-lm hash
curl -s "https://api.github.com/repos/ml-explore/mlx-lm/commits/main" | grep '"sha"' | head -1

# Download Python source
curl -s "https://raw.githubusercontent.com/ml-explore/mlx-lm/main/mlx_lm/models/cache.py" -o /tmp/cache.py

# Build and test
cd packages/swift && swift build -c release && swift test

# Regenerate models (to ensure compatibility)
pnpm hf2swift --model llama --output packages/swift/Sources/NodeMLXCore/generated/models/LlamaGenerated.swift
```
49 changes: 23 additions & 26 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -39,9 +39,9 @@ jobs:
uses: actions/cache@v4
with:
path: packages/swift/.build
key: swift-build-${{ runner.os }}-${{ hashFiles('packages/swift/Package.resolved', 'packages/swift/Package.swift') }}
key: swift-build-v3-${{ runner.os }}-${{ hashFiles('packages/swift/Package.resolved', 'packages/swift/Package.swift', 'packages/swift/Sources/**/*.swift') }}
restore-keys: |
swift-build-${{ runner.os }}-
swift-build-v3-${{ runner.os }}-

- name: Cache Xcode DerivedData
uses: actions/cache@v4
Expand Down Expand Up @@ -73,32 +73,29 @@ jobs:
env:
CODECOV_TOKEN: ${{ secrets.CODECOV_TOKEN }}

# Swift Tests with Coverage
- name: Run Swift tests with coverage
# Swift Tests (unit tests - no model downloads)
- name: Run Swift unit tests
working-directory: ./packages/swift
run: |
xcodebuild test \
-scheme NodeMLX \
-destination 'platform=macOS' \
-enableCodeCoverage YES \
-resultBundlePath ./test-results.xcresult \
2>&1 | xcbeautify || true

- name: Export Swift coverage
working-directory: ./packages/swift
run: |
# Convert xcresult to JSON format for Codecov
xcrun xccov view --report --json test-results.xcresult > coverage.json || true

- name: Upload Swift coverage to Codecov
uses: codecov/codecov-action@v5
with:
files: ./packages/swift/coverage.json
flags: swift
name: swift-coverage
fail_ci_if_error: false
env:
CODECOV_TOKEN: ${{ secrets.CODECOV_TOKEN }}
# Build tests with testing enabled (builds everything including library)
swift build -c release -Xswiftc -enable-testing --build-tests

# Copy Metal library to test bundle location
TEST_BUNDLE=".build/arm64-apple-macosx/release/NodeMLXPackageTests.xctest/Contents/MacOS"
METALLIB=".build/arm64-apple-macosx/release/mlx-swift_Cmlx.bundle/Contents/Resources/default.metallib"
if [ -f "$METALLIB" ]; then
mkdir -p "$TEST_BUNDLE"
cp "$METALLIB" "$TEST_BUNDLE/mlx.metallib"
echo "✓ Copied mlx.metallib to test bundle"
else
echo "⚠ mlx.metallib not found at $METALLIB - tests may fail"
find .build -name "*.metallib" 2>/dev/null || true
fi

# Run unit tests only (skip integration tests that require model downloads)
# Integration tests run via the smoke test below with cached models
# Filter pattern: list all unit test classes explicitly
swift test -c release --skip-build --filter 'GenerateTests|KVCacheTests|MaskTests|PerformanceTests|RoPEUtilsTests|SamplingUtilsTests|StringOrNumberTests|SwitchLayersTests'

- name: Verify Swift library
run: test -f packages/node-mlx/swift/libNodeMLX.dylib
Expand Down
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -47,3 +47,4 @@ DerivedData/
# Temporary files
*.tmp
*.bak
.venv/
Loading
Loading