Skip to content

When the staging file can't be created, pulls and pushes fail with 500 instead of bypassing the cache #57

Description

@matt-edmondson

What's wrong

Both object paths call store.OpenStaging(...) with no error handling:

  • Endpoints/ObjectRouteHandler.cs:332, the download miss path
  • Endpoints/ObjectRouteHandler.cs:226, the upload path

Inside, ObjectStore.OpenStaging runs Directory.CreateDirectory and FileStream.New(..., FileMode.CreateNew, ...) unguarded, so any I/O failure there escapes as a 500.

This breaks the project's own contract. StreamTee, ReadTeeStream and the README say that a failed store write only means the cache stays cold, and that this is "far better than a failed push". Once bytes are flowing that holds, but it does not hold when the staging file can't be opened in the first place. The writability probe runs only once at startup, so problems that appear later are not caught.

Failure scenario

The staging file can't be created, for example because of:

  • EACCES after a permissions change
  • a read-only remount after an ext4 error
  • inode exhaustion
  • a stray file at {root}/{upstream}/staging

On a download miss, the upstream response status and headers have already been copied when OpenStaging throws. The client gets a 500 even though upstream served the object. Every push fails the same way.

Reproduced by placing a file at {root}/github/staging in the integration fixture:

IOException: Cannot create '/gitlfscache-integration/github/staging' because a file or directory with the same name already exists.

There was 1 upstream fetch and the client received no object.

Suggested fix

Wrap OpenStaging in both handlers and catch IOException, UnauthorizedAccessException and ArgumentException. On failure:

  • download: log, record a store-failure metric, and continue with storeLocally: false, so the client still gets the bytes from upstream;
  • upload: relay the body to upstream without teeing it into the store.

Acceptance: with the staging directory unwritable, downloads and uploads through the cache still succeed, a warning is logged and a metric is incremented, and a test covers each path.

Activity

  1. matt-edmondson commented on Sep 28, 2026

    @matt-edmondson
    ContributorAuthor

    Triage

    • Category: Bug
    • Priority: High. When the store can't create a staging file (permissions change, read-only remount, inode exhaustion, a stray file), every pull miss and every push through the proxy fails with 500. That happens even though upstream is healthy and served the object. It breaks the project's own contract that a store failure only leaves the cache cold, and the startup probe can't catch conditions that appear later.
    • Area / suggested assignment: Objects, in GitLfsCache/Endpoints/ObjectRouteHandler.cs (lines 226 and 332) and ObjectStore.OpenStaging
    • Duplicates: none found among open issues in the org
    • In progress: no direct match. Open PR Refuse to publish a staging file whose write failed #59 hardens the staging write path (refuse to publish a failed write), which complements this open failure but doesn't fix it.

    Notes: Catch the failure around OpenStaging and fall back to storeLocally: false on download, and to a plain relay on upload. The warning and metric the acceptance criteria ask for make the degraded state visible without failing users.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

readyFully specified; implement as written

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions