Single-node S3-like object storage built with .NET 10. Supports streaming file operations, object versioning, multipart uploads, and PostgreSQL-based metadata storage.
English | Русский
SystemDesign.CloudStorage is a single-node object storage service inspired by the core concepts of Amazon S3.
The service separates logical objects from their physical representation. Users, buckets, object keys, versions, metadata, and multipart upload state are stored in PostgreSQL, while object contents are stored as binary blobs in an external filesystem.
The project intentionally focuses on a single storage node and does not implement replication, sharding, consistent hashing, distributed consensus, or cluster coordination.
flowchart LR
Client -->|JWT / HTTP Stream| API
API --> Application
Application --> Infrastructure
Infrastructure --> PostgreSQL[(PostgreSQL)]
Infrastructure --> FS[(File System)]
Cleanup[Background Cleanup] --> PostgreSQL
Cleanup --> FS
SystemDesign.CloudStorage.Api
SystemDesign.CloudStorage.Application
SystemDesign.CloudStorage.Domain
SystemDesign.CloudStorage.Infrastructure
SystemDesign.CloudStorage.Persistence
Api— REST API, Swagger, HTTP Range and conditional request handling.Application— contracts, DTOs, options, and application-level errors.Domain— users, buckets, logical objects, versions, blobs, and multipart models.Infrastructure— storage operations, JWT, filesystem access, concurrency control, quotas, and cleanup.Persistence— Entity Framework Core, PostgreSQL mappings, and migrations.
An object is identified by:
Owner + Bucket + Key
For example:
Bucket: documents
Key: users/123/passport.pdf
The object key is a logical identifier. It is never used directly as a filesystem path.
A physical blob may be stored as:
CloudStorage/
└── objects/
└── dc/
└── f9/
└── dcf982d10b5045808c8a1cf5bb73558c.blob
This keeps the logical namespace completely independent from physical storage.
erDiagram
USER ||--o{ REFRESH_TOKEN : owns
USER ||--o{ BUCKET : owns
BUCKET ||--o{ STORED_OBJECT : contains
STORED_OBJECT ||--o{ OBJECT_VERSION : versions
OBJECT_VERSION ||--o{ OBJECT_METADATA : has
OBJECT_VERSION }o--o| PHYSICAL_BLOB : references
BUCKET ||--o{ MULTIPART_UPLOAD : receives
MULTIPART_UPLOAD ||--o{ MULTIPART_PART : consists_of
PostgreSQL acts as the storage catalog. Actual object contents are never stored in the database.
Uploads and downloads are fully streamed.
HTTP Request Stream
│
▼
Temporary File
│
├── SHA-256 / ETag
▼
Physical Blob
│
▼
PostgreSQL Metadata
Objects are not loaded entirely into byte[] or MemoryStream. Memory consumption is therefore bounded by the streaming buffer rather than object size.
Uploads are first written to temporary storage. After the stream has been successfully written and validated, the temporary file is promoted to a physical blob and its metadata is committed to PostgreSQL.
Every PUT to an existing Bucket + Key creates a new immutable object version.
document.docx
│
├── Version 1 ──→ Blob 1
├── Version 2 ──→ Blob 2
└── Version 3 ──→ Blob 3
Each version has its own VersionId.
A regular GET returns the latest active version. A specific version can also be retrieved explicitly.
Deleting an object without specifying a version does not physically delete its previous versions.
Instead, a new delete-marker version is created:
StoredObject
│
├── Version 1 ──→ PhysicalBlob
└── Version 2 ──→ DeleteMarker
The regular object lookup then returns 404, while previous versions and their blobs remain available through their VersionId.
Large objects can be uploaded in multiple independent parts:
Initiate
│
├── Upload Part 1
├── Upload Part 2
├── Upload Part 3
│
▼
Complete
Each part is streamed into temporary storage and receives its own integrity information.
During completion, parts are validated and sequentially combined into the final blob without loading the entire object into memory.
Multipart upload is optional. Regular streaming PUT also supports large objects.
Downloads support HTTP byte ranges:
Range: bytes=1048576-2097151The service seeks directly to the requested position in the physical blob and returns only the requested range with 206 Partial Content.
Also supported:
HEADETagIf-MatchIf-None-MatchContent-RangeAccept-Ranges
Bucket contents can be queried using:
- prefix filtering;
- page size;
- continuation tokens;
- keyset pagination.
Filtering and pagination are performed by PostgreSQL rather than loading the entire bucket into application memory.
The API uses JWT Bearer authentication.
Supported operations include:
Register
Login
Refresh
Revoke
Users can authenticate using either login/password or email/password.
Access tokens have a one-hour lifetime and are not persisted. Refresh tokens are cryptographically generated, while only their SHA-256 hashes are stored in PostgreSQL.
Concurrent writes to the same logical object are serialized using PostgreSQL transaction-level advisory locks based on:
Owner + Bucket + Key
Writes to unrelated object keys do not require a global application lock.
Filesystem operations and PostgreSQL transactions cannot participate in the same atomic transaction. The storage therefore uses temporary files, atomic promotion, compensating deletion, and background orphan cleanup to maintain consistency.
A background service periodically removes:
- abandoned multipart uploads;
- expired multipart parts;
- stale temporary files;
- orphan physical blobs that are no longer referenced.
Delete markers do not cause previous object versions to be removed.
Configurable limits include:
MaxObjectSize
MaxMultipartPartSize
MinMultipartPartSize
MaxParts
MaxBucketsPerUser
MaxStorageBytesPerUser
Physical object storage must be located outside the application and repository directories.
Example for Windows:
{
"Storage": {
"RootPath": "D:\\CloudStorage"
}
}Example for Linux:
/var/lib/systemdesign-cloudstorage
The service creates its internal storage directories automatically:
CloudStorage/
├── objects/
├── temporary/
└── multipart/
Requirements:
- .NET 10 SDK
- PostgreSQL
- configured storage root
- JWT signing key
SystemDesign.CloudStorage is intentionally designed for a single storage node.
The following distributed storage mechanisms are outside its current scope:
- replication;
- sharding;
- consistent hashing;
- erasure coding;
- cluster membership;
- distributed consensus;
- automatic rebalancing.
Однонодовое S3-подобное объектное хранилище на .NET 10. Поддерживает потоковую работу с файлами, версионирование, multipart upload и хранение метаданных в PostgreSQL.
English | Русский
SystemDesign.CloudStorage — однонодовое объектное хранилище, реализующее основные принципы S3.
Сервис разделяет логическое представление объектов и их физическое содержимое. Пользователи, buckets, ключи объектов, версии, метаданные и состояние multipart uploads хранятся в PostgreSQL, а содержимое объектов — в виде бинарных blobs во внешней файловой системе.
Архитектура намеренно рассчитана на одну storage-ноду и не включает репликацию, шардирование, consistent hashing, distributed consensus и координацию кластера.
flowchart LR
Client[Клиент] -->|JWT / HTTP Stream| API
API --> Application
Application --> Infrastructure
Infrastructure --> PostgreSQL[(PostgreSQL)]
Infrastructure --> FS[(Файловая система)]
Cleanup[Фоновая очистка] --> PostgreSQL
Cleanup --> FS
SystemDesign.CloudStorage.Api
SystemDesign.CloudStorage.Application
SystemDesign.CloudStorage.Domain
SystemDesign.CloudStorage.Infrastructure
SystemDesign.CloudStorage.Persistence
Api— REST API, Swagger, HTTP Range и conditional requests.Application— контракты, DTO, options и прикладные ошибки.Domain— пользователи, buckets, логические объекты, версии, blobs и multipart-модели.Infrastructure— операции хранилища, JWT, файловая система, concurrency control, quotas и cleanup.Persistence— Entity Framework Core, mappings PostgreSQL и migrations.
Объект идентифицируется комбинацией:
Owner + Bucket + Key
Например:
Bucket: documents
Key: users/123/passport.pdf
Key является логическим идентификатором и никогда напрямую не используется как путь файловой системы.
Физически blob может находиться:
CloudStorage/
└── objects/
└── dc/
└── f9/
└── dcf982d10b5045808c8a1cf5bb73558c.blob
Таким образом, логическое имя объекта полностью отделено от его физического расположения.
erDiagram
USER ||--o{ REFRESH_TOKEN : owns
USER ||--o{ BUCKET : owns
BUCKET ||--o{ STORED_OBJECT : contains
STORED_OBJECT ||--o{ OBJECT_VERSION : versions
OBJECT_VERSION ||--o{ OBJECT_METADATA : has
OBJECT_VERSION }o--o| PHYSICAL_BLOB : references
BUCKET ||--o{ MULTIPART_UPLOAD : receives
MULTIPART_UPLOAD ||--o{ MULTIPART_PART : consists_of
PostgreSQL выполняет роль каталога object storage. Само содержимое объектов в базе данных не хранится.
Upload и download выполняются потоково.
HTTP Request Stream
│
▼
Temporary File
│
├── SHA-256 / ETag
▼
Physical Blob
│
▼
PostgreSQL Metadata
Объект не загружается целиком в byte[] или MemoryStream. Потребление памяти ограничивается streaming buffer и практически не зависит от размера файла.
При загрузке данные сначала записываются во временный файл. После успешного завершения записи файл переносится в постоянное blob storage, а информация об объекте и его версии сохраняется в PostgreSQL.
Каждый повторный PUT существующего Bucket + Key создаёт новую immutable-версию:
document.docx
│
├── Version 1 ──→ Blob 1
├── Version 2 ──→ Blob 2
└── Version 3 ──→ Blob 3
Каждая версия получает собственный VersionId.
Обычный GET возвращает последнюю активную версию. Конкретную версию можно получить отдельно по её VersionId.
Удаление объекта без указания версии не удаляет предыдущие версии физически.
Вместо этого создаётся новая версия с delete marker:
StoredObject
│
├── Version 1 ──→ PhysicalBlob
└── Version 2 ──→ DeleteMarker
После этого обычный GET возвращает 404, но предыдущие версии и соответствующие blobs продолжают существовать и доступны по VersionId.
Крупные объекты могут загружаться отдельными частями:
Initiate
│
├── Upload Part 1
├── Upload Part 2
├── Upload Part 3
│
▼
Complete
Каждая часть потоково сохраняется во временное хранилище и получает информацию для проверки целостности.
При Complete части проверяются и последовательно объединяются в final blob без загрузки всего объекта в память.
Multipart не является обязательным: обычный streaming PUT также поддерживает крупные файлы.
Download поддерживает HTTP Range:
Range: bytes=1048576-2097151Сервис выполняет seek непосредственно внутри физического blob и возвращает только запрошенный диапазон с 206 Partial Content.
Также поддерживаются:
HEAD;ETag;If-Match;If-None-Match;Content-Range;Accept-Ranges.
Для просмотра содержимого bucket поддерживаются:
- prefix filtering;
- page size;
- continuation token;
- keyset pagination.
Фильтрация и pagination выполняются непосредственно PostgreSQL без загрузки всего списка объектов в память приложения.
API использует JWT Bearer Authentication.
Поддерживаются:
Register
Login
Refresh
Revoke
Вход возможен по login/password или email/password.
Access token имеет lifetime один час и не сохраняется в БД. Refresh token генерируется криптографически безопасным способом, а в PostgreSQL сохраняется только его SHA-256 hash.
Конкурентные записи одного logical object сериализуются с помощью PostgreSQL transaction-level advisory locks на основании:
Owner + Bucket + Key
Операции над разными keys не требуют глобальной блокировки приложения.
PostgreSQL и файловая система не могут участвовать в общей атомарной транзакции. Для поддержания согласованности используются temporary files, atomic promotion, compensating deletion и фоновая очистка orphan blobs.
BackgroundService периодически очищает:
- abandoned multipart uploads;
- expired multipart parts;
- старые temporary files;
- orphan blobs, на которые больше не ссылаются версии объектов.
Наличие delete marker само по себе не приводит к удалению предыдущих версий объекта.
Поддерживаются конфигурируемые лимиты:
MaxObjectSize
MaxMultipartPartSize
MinMultipartPartSize
MaxParts
MaxBucketsPerUser
MaxStorageBytesPerUser
Физическое хранилище должно находиться вне каталога приложения и repository.
Пример для Windows:
{
"Storage": {
"RootPath": "D:\\CloudStorage"
}
}Для Linux:
/var/lib/systemdesign-cloudstorage
Сервис автоматически создаёт внутренние директории:
CloudStorage/
├── objects/
├── temporary/
└── multipart/
Необходимы:
- .NET 10 SDK;
- PostgreSQL;
- настроенный
Storage:RootPath; - JWT signing key.
SystemDesign.CloudStorage рассчитан на одну storage-ноду.
Намеренно не реализованы:
- replication;
- sharding;
- consistent hashing;
- erasure coding;
- cluster membership;
- distributed consensus;
- automatic rebalancing.
Основное внимание уделено механизмам object storage внутри одной ноды: streaming I/O, versioning, physical blob storage, metadata, consistency, multipart uploads и lifecycle management.