Garuda is a Swift web framework and the HTTP server underneath it. It is my attempt to create a spiritual successor to Kitura: a server written in Swift, with its own HTTP engine, for the protocols and the application features a production service needs today.
The application API is modelled after axum. Routes, typed extractors, and the set of features a production framework is expected to offer follow axum's shape. Garuda and axum compares the two area by area. The HTTP engine is Garuda's own.
The same route in each. axum names the parameter in braces; Garuda uses a colon. Both take it as a typed Path and answer JSON.
// axum
async fn person(Path(id): Path<i32>) -> Json<Person> {
Json(Person { id, name: "Ada".into() })
}
let app = Router::new().route("/person/{id}", get(person));// Garuda
app.get("/person/:id") { (id: Path<Int>) async in
JSON(Person(id: id.value, name: "Ada"))
}The key difference is where the handler runs. axum makes every handler a future and runs it on Tokio, a work-stealing scheduler: an idle thread takes a waiting task from a busy one. SwiftNIO is that design in Swift. An event loop owns the socket, and the handler is a task the runtime resumes.
Garuda does not use either. The worker thread that read the request runs the handler, and a response that can be sent at once is written on that thread. There is no hop onto a scheduler before those bytes go out, which is what SwiftNIO would insert between the read and the write. The engine parses HTTP/1.1, HTTP/2 and HTTP/3 itself.
A worker is one process and one thread. Nothing is shared across workers, so nothing is locked, and a crash takes only that worker's connections. That boundary matters more in Swift than in Rust: a force-unwrapped nil ends the process, where Tokio catches a panic inside the task that raised it. Processes cannot steal tasks from each other. A worker that is ahead leaves new connections to the others, and can hand idle HTTP/1 keep-alive connections to a worker where they would not wait (--balance). One process per worker has the rest.
Swift 6.2 or newer. No Foundation, no SwiftNIO.
This project is an independent effort. It is not an IBM product, and it is not a continuation maintained by the Kitura project.
Released: 1.0. The surface below is implemented and tested. COMPATIBILITY.md says what a release may change, and Status below lists what is missing.
| Area | What Garuda provides |
|---|---|
| HTTP | HTTP/1.1 with a parser strict where that prevents request smuggling, HTTP/2 over TLS and in cleartext, and HTTP/3 over QUIC |
| TLS and certificates | TLS 1.2 and 1.3, several certificates chosen by SNI, and ACME renewal with tls-alpn-01 on the port already served |
| Handlers | Routes, groups, and middleware. A synchronous handler is a direct call on the worker thread. An async handler runs on a task reused from that worker's pool |
| Typed API | Extractors for path, query, JSON body, form, and multipart. Answers as JSON, HTML, text, bytes, or a redirect. A type can validate its values; every broken rule is returned together as 422 |
| OpenAPI | An OpenAPI 3.1 document generated from the routes, and Swagger UI |
| Application middleware | Deadlines, CORS, bearer and basic authentication, authorization policies, CSRF protection, security headers, signed and encrypted cookies, and sessions in memory, Redis, or SQLite |
| Identity | JWT (HS, RS, PS, ES, and EdDSA), JWKS verification, rotating refresh tokens, and PBKDF2-HMAC-SHA256 password hashing |
| HTTP client | HTTP/1.1 and HTTP/2, redirects by policy, decompression, streamed bodies, and server-sent events |
| Data | PostgreSQL and Redis on the worker's poller, including Redis Cluster and Sentinel, and SQLite through the system's libsqlite3 |
| Streaming | Streamed responses and request bodies, server-sent events, broadcast across worker processes, and interim 1xx responses such as 103 Early Hints |
| WebSockets | WebSockets over HTTP/1.1, HTTP/2, and HTTP/3. The engine joins fragments, checks UTF-8, answers pings, and supports permessage-deflate |
| WebTransport | WebTransport over HTTP/3 on the same port and routes, with bidirectional and unidirectional streams and datagrams |
| IETF Resumable Uploads | The IETF resumable upload protocol (draft-ietf-httpbis-resumable-upload, interop version 9). A client cut off mid-upload resumes on any worker and over any of the three HTTP protocols |
| Files and compression | Static files with sendfile, ETag, and byte ranges, including precompressed copies. Response and request compression with brotli, zstd, and gzip |
| Request limits | A body limit and a concurrency limit on a group of routes. Rate limiting counted across workers |
| Edge policy | HTTPS redirect, HSTS, trusted proxy headers, allowed hosts, and client address allow and deny lists |
| Operations | Prometheus metrics, health checks, request IDs, W3C trace context, an access log, a shared response cache, graceful shutdown, and reload on SIGHUP. Unix sockets and more than one worker |
| Tracing | Spans through swift-distributed-tracing for the request, outbound HTTP, PostgreSQL, and Redis |
| Tests | app.test runs the real engine in the test process, over a socket pair |
| Process model | A supervisor accepts connections and forks one process per worker. In-memory state is per worker |
| Platforms | Linux and macOS 15, on Swift 6.2 or newer |
Outside this release: QUIC 0-RTT and QUIC session resumption, kernel TLS, multipart/byteranges, reads from Redis replicas, and Windows other than through WSL 2.
Every server runs the same three-route application and the same load command, each in an invocation of its own, and the six rotate so that no server is always measured last. Garuda's range and axum's overlap, so these rounds do not separate them; ntex is ahead of both, and Garuda is 2.8x Hummingbird and 3.6x Vapor. BENCHMARKS.md has the method, the rounds behind each bar, and the runs where Garuda is slower.
External benchmark: the-benchmarker measures Garuda on its own hardware and load command. The chart above is a separate run on our hardware.
swift build -c release
.build/release/garuda --port 8000 --workers 0
curl -i http://127.0.0.1:8000/user/42The garuda binary serves a small benchmark application. Your own application
depends on the Garuda library:
// Package.swift
dependencies: [.package(url: "https://github.com/grepjava/garuda", from: "1.1.0")],
targets: [.executableTarget(name: "app", dependencies: [.product(name: "Garuda", package: "garuda")])]// Sources/app/main.swift
import Glibc // Darwin on macOS
import Garuda
struct Person: Codable { let id: Int; let name: String }
let app = Application()
app.get("/person/:id") { (id: Path<Int>) async in
JSON(Person(id: id.value, name: "Ada"))
}
exit(app.run())app.run() reads the same command-line flags as the garuda binary. Swift 6.2
or newer, on Linux or macOS 15 -- 6.2 because a borrowed body is a Span. INSTALLATION.md lists the
system packages, certificates and running as a service.
Examples/ has complete applications to run and copy from: a CRUD API on SQLite, accounts with hashed passwords and sessions, streaming both ways, chat rooms over WebSockets across workers, a file service with resumable uploads and ranged downloads, per-user photo uploads that are authenticated and resumable, a network probe over WebTransport, and a whole starter application on PostgreSQL.
Routes are registered with get, head, post, put, delete, patch, options, or on. A path segment is a literal, a parameter (:param), or a trailing rest segment (*rest). HEAD falls back to GET, and a known path under the wrong method is 405 with Allow.
A handler declares what it needs and returns what it means:
app.get("/search") { (query: Query<Search>) async in JSON(results(for: query.value)) }
app.post("/people") { (body: Body<NewPerson>) async in JSON(create(body.value), status: .created) }
app.post("/login") { (form: Form<Credentials>) async in Redirect(to: "/") }- Extractors:
Path<T>(the next path parameter, percent-decoded),Query<T>,Body<T>(JSON),Form<T>,Multipart,State<T>,Context<Key>, or your ownRequestExtractor. AnAsyncRequestExtractormay await, to load the signed-in user from a database, for an async handler.E?is nil whereEwould refuse, andResult<E, any Error>hands the handler the refusal. A value that will not decode is a 400 that says what was wrong. - Answers:
JSON,HTML,Text,Bytes,Redirect, aString, anHTTPStatus, or an Optional whosenilis a 404. - Errors: a thrown
ResponseErroris the response.throw HTTPError(.conflict, "the name is taken")answers 409 with{"error":"the name is taken"}. Anything else thrown is a 500 and a log line. - Rules: a type saying what its values have to be, beyond their shape, conforms to
Validated, and whatever is decoded from a request is held to them before the handler runs.
struct NewOrder: Decodable, Validated {
let quantity: Int
let email: String
func validate(_ check: inout Validation) {
check.range("quantity", quantity, atLeast: 1, atMost: 100)
check.email("email", email)
}
}Every broken rule is answered at once, as 422 with the fields named: {"error":"quantity must be at least 1","fields":[{"field":"quantity","message":"must be at least 1"}]}. A decoding failure fills fields the same way, from the path it already reports, so one answer tells a form where each message goes whichever of the two refused the request.
JSON goes through Garuda's own coder over Encodable and Decodable, which
reads only the keys a type asks for, straight from the request bytes. @JSON
takes a type off Codable altogether:
import GarudaJSON
@JSON struct Order: Codable {
let id: Int
let name: String
let tags: [String]
}Nothing at the call site changes -- the same Body<Order>, the same
JSON(order) -- and the type keeps Codable for everything else.
@PostgresRow, from GarudaSQL, does the same for a database row: a
property is read from the column of its own name, and first(Item.self, ...)
is written the same way either way. The macro
writes the reading and the writing out longhand; the bytes are the same ones
Codable sent, and the json workload's request went from 15.6 microseconds to
13.2 on one worker. It lives in its own module because it is the only part of
Garuda that needs swift-syntax.
Under the typed API, a handler can take the request and response directly:
app.onAsync(.get, "/user/:id") { request, response in
request.withParameter(0) { response.send($0) } // lent bytes, no copy
}Request and Response are ~Copyable. The request lends its bytes to a
closure as a Span, which the compiler keeps from escaping (withPath,
withHeader, withBody and others). path, header(_:), body and the rest
return owned copies.
A closure that awaits is an async handler. It runs on a task reused from the worker's pool, on the worker's own thread, and allocates nothing per request once warm.
app.deadline(milliseconds: 500) {
app.onAsync(.get, "/weather/:city") { request, response in
let client = request.client // read before the first await
let city = request.parameter(0)
let answer = try await client.get("https://api.example.com/\(city)")
response.send(bytes: answer.body, contentType: "application/json")
}
}A request still unanswered at its deadline is answered 504. A closed connection
or reset stream cancels a handler waiting on the engine. request.client
speaks HTTP/1.1, and HTTP/2 over TLS when the server offers it. It resolves
names on the worker's poller and keeps connections for reuse. It asks for
gzip, deflate, brotli and zstd and decodes what comes back, and it follows
redirects only when client.redirects says to: .sameOrigin(), .any() or
.matching { url in … }. client.timeoutMilliseconds bounds each wait;
client.totalTimeoutMilliseconds bounds the whole exchange, redirects included,
for work no route deadline covers.
client.stream returns once the response head is in, and the body is read as
it arrives -- a large file relayed on, an upstream's server-sent events, a
model's tokens -- rather than held whole:
app.onAsync(.get, "/relay") { request, response in
let client = request.client // read before the first await
let upstream = try await client.stream(.get, "https://llm.example/v1/stream")
let body = response.stream(contentType: "text/event-stream")
while let event = try await upstream.nextEvent() {
try await body.write("data: \(event.data)\n\n")
}
}A streamed body is not held to maxBodyBytes. A caller that stops reading
stops the upstream: over HTTP/1.1 by TCP's window, over HTTP/2 because the
stream's window opens only as the caller reads. cancel() gives up the rest.
app.state { _ in PostgresPool(PostgresConfiguration(host: "db", user: "app", password: secret)) }
app.group("/api") {
app.cors(CORSPolicy(origins: ["https://app.example.com"], allowCredentials: true))
app.authenticate(bearer: CurrentUser.self) { token in try await sessions.user(token) }
app.use { request, response async throws -> (any ResponseConvertible)? in
response.onSend { outgoing in outgoing.addHeader("cache-control", "no-store") }
return nil
}
app.get("/user/:id") { (id: Path<Int>, db: State<PostgresPool>) async throws in
try await db.value.first(User.self, "select id, name from users where id = $1", id.value).map { JSON($0) }
}
}app.statebuilds a value once in each worker process, after the fork. One value per type; registering a type twice is refused.app.problems()is what the routes cannot do, found before serving: a handler taking morePathextractors than its pattern has parameters, or asking for aState<T>nothing registered.run()prints them and exits 2 instead of starting, and a test can assertapp.problems().isEmpty.app.useruns before every route in its scope, whether it is called before or after the routes. It returnsnilto carry on or an answer to send instead, and can be async.response.onSendsees and changes the final response, whoever answered.request[context: Key.self]carries values from middleware to the handler, which reads them withContext<Key>.app.corssets the scope's CORS policy. It answers preflights, including to paths routed only for other methods, and runs ahead of the scope's middleware, so a preflight never meets authentication and a 401 still reaches the page.app.authenticate(bearer:)andapp.authenticate(basic:)read the Authorization header and keep who it belongs to in the context, or answer 401 with the challenge. Withstate:, the check is handed whatapp.statebuilt, such as the database sessions live in.BearerTokenandBasicCredentialsare the same as extractors, andconstantTimeEqualscompares secrets.app.authorize(CurrentUser.self, .admin)requires a named rule of every route in the scope, onceauthenticatehas said whose request it is. APolicyis(Value) -> Boolwith a name, combined withand,orandabout, and callable in a test without a request. A rule that does not hold is 403 saying what would have been enough; nobody under the key is 401.authorize(jwt:)reads the claims a scope verified, andPolicy.scope("orders:write")an OAuth 2.0 scope in them. Bothauthenticateandauthorizeadd what they can answer, and the security scheme, to every route of the scope in the OpenAPI document.JWT<Claims>takes a verified JSON Web Token (HS, RS, PS, ES or EdDSA), keys from PEM, JWK or a secret, withexp,nbf,issandaudchecked;keys.sign(claims)issues one andkeys.publicJWKSpublishes the key set.JWKSVerifierchecks tokens from an identity provider against its JWK Set.TokenIssuerpairs short-lived access tokens with refresh tokens that rotate on every use and revoke their whole chain when a spent one is reused.Passwords.hashandPasswords.verifyuse PBKDF2-HMAC-SHA256 on the blocking pool.Tokens.random()makes a session token andTokens.digestwhat to store in its place.request.log.info("order placed", ["order": "\(id)"])writes to the application log with the request's method, path, request ID and trace context on the line.AppLogis the same log outside a request.--log-format jsonmakes every line one JSON object.app.onResponse { done in … }sees every request once it is answered: the matched route's pattern, status, duration, request ID and trace, and why a route failed. This is what per-route metrics or error reporting hang from.app.maxBodySize(64 << 20) { … }gives a scope's routes a body limit of their own in place of--max-body, applied before the body is read.app.concurrencyLimit(8) { … }lets that many of a scope's handlers run at once in each worker, and answers the next 503.request.cookie("theme")andresponse.setCookie(Cookie("theme", "dark"))read and set cookies. A cookie isPath=/,HttpOnlyandSameSite=Laxunless it says otherwise, andSecureover HTTPS. With aCookieKey, a cookie is signed (HMAC-SHA256) or encrypted (AES-256-GCM), and one that was tampered with reads as absent. The key takes previous secrets for rotation.app.sessions(store: …)gives a scope's routes aSession: a map of strings kept in memory, Redis or SQLite, found by a random ID in a cookie.try await session.set("user", id)writes it to the store before it returns.session.renew()moves it to a new ID at login, andsession.destroy()deletes it and its cookie.app.csrfProtection()refuses, with 403, a POST, PUT, PATCH or DELETE that the browser says another site's page started, from Sec-Fetch-Site or else Origin against Host. It needs no tokens in forms.trustedOriginslets a front end on another origin through.app.securityHeaders()puts nosniff, frame, referrer and cross-origin policies on every answer in its scope, and Strict-Transport-Security on those over HTTPS. A Content-Security-Policy is one field away, and a header a route sets itself is kept.app.trailingSlash(.redirect)answers/users/with a 308 to/userswhen only that is routed, and.ignoreserves it from/usersdirectly. Routes match exactly by default.app.requestDecompression()decodes a gzip, deflate, br or zstd request body before the handler reads it, held to the route's body limit as it inflates.app.allowedHosts(["example.com", "*.example.com"])answers 400 to a request for any other Host, andapp.addressFilter(allow: ["10.0.0.0/8"])answers 403 to a client address outside the list or on adenylist.
app.tracing() traces through
swift-distributed-tracing,
so any tracer written for it records the spans, the OpenTelemetry exporters
among them:
import Tracing
app.onWorkerStart { _ in
InstrumentationSystem.bootstrap(makeTracer()) // any Tracer
}
app.tracing()Each request is a server span named for its route, GET /users/:id,
continuing the trace its headers carry. Under it are child spans for every
call request.client makes, which passes the trace on in its headers, and for
every PostgreSQL statement and Redis round trip. Spans a handler starts with
withSpan go under it too. The attributes follow OpenTelemetry's HTTP and
database conventions. SQL is recorded as written; bound values, Redis keys and
values, and query-string values in URLs are not. A streamed response's span
ends with its last byte, not its head.
Workers are processes, and an exporter's threads do not survive the fork, so
the tracer is made in each worker: bootstrap it in onWorkerStart, or pass
app.tracing { worker in … } a closure that returns one. With tracing off,
each request pays one check.
app.openAPI(OpenAPIInfo(title: "Shop", version: "1.0.0")) serves an OpenAPI
3.1 document at /openapi.json, and app.swaggerUI() serves Swagger UI at
/docs. The document comes from the routes: a typed route's extractors and
return type give its parameters, request body, security and response, with
JSON Schemas read from the Decodable types, so nothing needs annotating.
Every route method returns an OpenAPIOperation for what types cannot say:
app.get("/orders/:id") { (id: Path<Int>) async throws in JSON(try await order(id.value)) }
.summary("An order by its number")
.tags("orders")
.response(.notFound, "No order has that number")A native driver runs on the worker's poller. It supports SCRAM-SHA-256, TLS
(required by default) and a pool per worker with an acquire timeout. Rows decode
into Decodable types by column name. db.transaction { tx in … } commits or
rolls back. Statements stay prepared on each connection, and values are read in
binary after the first run. UUID, Timestamp, PostgresDate,
PostgresTime, PostgresInterval, PostgresNumeric (exact, kept as its
digits) and PostgresJSON<T> come without Foundation. An array column is a
Swift list -- [String] for text[] -- and a list binds as one. app.listen
hears NOTIFY in each worker, on a connection of its own, reconnecting with
its channels when the server restarts. copyIn and copyOut are COPY in
both directions, a chunk at a time; a composite type is a PostgresRecord and
an enum is a Swift enum; and a server on this machine is reached over its unix
socket.
app.state { _ in RedisPool(RedisConfiguration(host: "cache", password: secret)) }
app.get("/visits/:page") { (page: Path<String>, redis: State<RedisPool>) async throws in
String(try await redis.value.incr("visits:\(page.value)"))
}A native driver runs on the worker's poller, speaking RESP3 through HELLO and
RESP2 to servers older than Redis 6; Valkey works the same. TLS is required by
default, and ACL users, databases and unix sockets are supported. Commands come
with typed replies (get, set with expiry and NX/XX, hashes, lists, sets,
getJSON), and send takes any other. pipeline and transaction go in one
round trip, session holds one connection for WATCH, and subscribe listens
to channels and patterns on a connection of its own.
RedisCluster is the same API across a cluster: a pool per node, a slot map
learned from the cluster, and every command aimed at the node that owns its
key -- following MOVED and ASK when the map is behind. RedisSentinelPool
asks a set of sentinels where the master is, checks what they name with ROLE,
and asks again when a failover takes it away. Garuda's session and
refresh-token stores work on either.
app.state { _ in
let db = try SQLiteDatabase(SQLiteConfiguration(path: "/var/lib/app/app.db"))
try db.migrate(["create table notes (id integer primary key, text text not null)"])
return db
}
app.get("/note/:id") { (id: Path<Int>, db: State<SQLiteDatabase>) async throws in
try await db.value.first(Note.self, "select id, text from notes where id = ?", id.value).map { JSON($0) }
}The system's libsqlite3 is loaded at run time, and every statement runs on the
worker's blocking pool, so a worker keeps serving while SQLite reads the disk
or waits for another process's lock. Each worker has one connection that
writes and up to four that read, in write-ahead-log mode; a statement moves to
a reader once SQLite has said it cannot write. Rows decode into Decodable
types, transaction begins IMMEDIATE, and migrate brings the schema up to
date by user_version when each worker starts.
app.get("/export") { () async in
StreamingBody(contentType: "text/csv") { body in
for row in rows { try await body.write(row.csv) } // waits while the client is behind
}
}
app.get("/ticks") { () async in
EventStream { events in
for i in 0... { try await events.send("tick \(i)"); try await events.sleep(milliseconds: 1000) }
}
}
app.webSocket("/chat/:room") { (ws: WebSocket, room: Path<String>) async throws in
for try await message in ws { try await ws.send(message) } // whole messages
}
struct Said: Codable { var text: String }
let lobby = Topic("lobby")
app.post("/say") { (said: Body<Said>) async in
try lobby.publish(said.value.text, event: "said") // heard by subscribers on every worker
return HTTPStatus.noContent
}
app.get("/lobby") { (last: LastEventID) async in
EventStream { events in try await events.forward(lobby, after: last) } // replays what a reconnect missed
}
app.webTransport("/room/:id") { (session: WebTransportSession, id: Path<Int>) async throws in
while let stream = try await session.acceptStream() { /* read and write */ }
}A streamed body is chunked on HTTP/1.1 and DATA frames on HTTP/2 and HTTP/3.
A write waits while more than 512 KiB is queued (writeHighWaterMark), and a throw part-way resets the
stream rather than ending it cleanly. A WebSocket handler sees whole messages:
the engine joins fragments, checks UTF-8, answers pings, sends keepalive pings,
runs the close handshake and, with --ws-compress, permessage-deflate. Messages
the handler has not read wait in a bounded queue, and past it the socket is not
read, so a fast sender is slowed. A send on a socket whose connection has
ended throws WebSocketError.closed, so a room holding sockets learns who has
gone rather than writing to them. The same route serves WebSockets over
HTTP/1.1, HTTP/2 and HTTP/3; on the last two, flow control slows just that
stream. WebTransport runs over HTTP/3 on the same
port and routes, with bidirectional and unidirectional streams, datagrams and
close codes. The webtransport example
uses all of them, with a page that runs it from a browser.
app.onStreamingBody(.put, "/files/:name", maxBodySize: 10 << 30) { request, response, body in
while let bytes = try await body.read() { try file.write(bytes) } // as it arrives
response.send(status: .created)
}
import GarudaUploads
app.resumableUploads("/uploads", store: try FileUploadStore(directory: "/var/lib/app/uploads"),
limits: UploadLimits(maxSize: 10 << 30)) { upload in
try moveIntoPlace(upload.path)
try upload.remove()
return HTTPStatus.created
}A route registered with onStreamingBody runs as soon as the request's head
arrives, and reads the body as it comes, with its own size limit. Reading is
flow control: bytes the handler has not read hold back the client over
HTTP/1.1, HTTP/2 and HTTP/3, so a slow disk slows only its own upload. If the
client goes away part-way, the handler still reads everything that arrived and
then gets RequestBodyError.incomplete.
GarudaUploads implements the IETF resumable upload protocol
(draft-ietf-httpbis-resumable-upload, interop version 9). A client cut off
mid-upload asks how much arrived and sends the rest, over any protocol and on
any worker. response.sendInterim sends 1xx responses such as 103 Early Hints.
The request that creates an upload and the request that completes it are
different requests -- often minutes apart, and on any worker -- so onCreate
is given the creating one, already through whatever middleware guards the
routes, and what it returns is kept on the upload for the completion handler:
app.authenticate(jwt: Claims.self)
app.resumableUploads("/photos", store: store, onCreate: { request in
var parameter = 0
let jwt = try JWT<Claims>.extract(from: request, parameter: ¶meter)
return ["user": jwt.claims.sub] // server-set metadata
}) { upload in
try record(upload.path, owner: upload.metadata["user"])
return HTTPStatus.created
}The server writes that metadata and no client can reach it, which is what makes
it the place for whose upload this is -- Content-Type and
Content-Disposition come from the client and are not. A throw from onCreate
refuses the upload before anything is written, so a request turned away leaves
nothing behind to expire.
uploads is relative to wherever the routes are mounted, so a group carries
them whole: inside group("/api") the uploads live at /api/uploads/:id and
that is what the Location says, and inside group("/users/:user") each
upload's URL sits under the user whose request created it.
The answer to the request that completes an upload is kept, Location
included, and a GET of the upload's URL gives it again until maxAge: a
client whose connection died as the upload finished still learns what became
of it. An answer whose body is streamed is not kept, since its bytes are
written after it has gone; return one whole if it should survive. The
uploads example puts
this together: per-user photos, owned through onCreate, checked before
they are kept, and resumable under a group.
A 104 is an interim response, and an intermediary is free to drop one: a CDN or
a reverse proxy in front of the server may give the client only the final 201.
A client should read the upload's URL from the 104 when it arrives and from the
Location of the 201 when it does not, which is what a client creating an
upload with an empty body does anyway.
@Test func personIsFound() throws {
let response = try app.test.get("/person/1")
#expect(response.status == .ok)
#expect(try response.json(Person.self).name == "Ada")
}app.test runs the real engine in the test process over a socket pair, with no
port: parsing, routing, middleware, handlers, timers and the response path.
The application API is modelled after axum. axum is a mature, widely used framework, so this is where Garuda stands against it, area by area.
| Area | axum (Tokio) | Garuda |
|---|---|---|
| Route dispatch | Every handler is a future the runtime polls | A synchronous handler is a direct call on the worker thread |
| Async handlers | async fn on a work-stealing pool |
Reused tasks on the worker's own executor, no allocation per request |
| Spreading load over cores | Idle threads steal tasks from busy ones | A worker ahead of the others leaves new connections to them; quick connections stuck behind slow requests move to a worker where they would not wait (differs) |
| Typed extraction | Path, Query, Json, Form, Multipart |
Path, Query, Body, Form, Multipart |
| Custom extractors | FromRequestParts, FromRequest, Option<T>, Result<T, E> |
RequestExtractor, AsyncRequestExtractor, E?, Result<E, any Error> |
| OpenAPI | utoipa or aide, with derive macros | app.openAPI, app.swaggerUI; schemas read from Decodable types |
| State | State<T>, one Arc shared by every thread |
State<T>, built in each worker process (differs) |
| Request-scoped values | Extension<T> |
request[context:], Context<Key> |
| Errors as responses | IntoResponse |
ResponseError, HTTPError |
| Protocols | HTTP/1.1 and HTTP/2 through hyper | HTTP/1.1, HTTP/2 and HTTP/3 |
| TLS | rustls or OpenSSL via axum-server; ACME from another crate |
Built in, with ACME |
| Nesting and 405 | nest, merge, fallback, 405 with Allow |
group, Router with nest and merge, fallback per scope, 405 with Allow |
| Middleware | Tower layers that wrap the handler | use before the handler; onSend on the response (differs) |
| Ready-made middleware | tower-http, tower-sessions, axum-extra | Server flags for compression, rate limits, request IDs, trace context, access log; app.deadline, app.cors, app.authenticate, JWT<Claims>, request.log, app.onResponse, app.maxBodySize, app.concurrencyLimit, cookies, app.sessions, app.csrfProtection, app.securityHeaders, app.trailingSlash, app.requestDecompression, app.allowedHosts, app.addressFilter |
| Streaming responses | Body::from_stream |
response.stream(), StreamingBody, with backpressure |
| Server-sent events | Sse, with keep-alive |
EventStream, with keep-alive comments and Last-Event-ID |
| Broadcast | tokio::sync::broadcast, within one process |
Topic, across worker processes, to event streams, WebSockets and long polls, with replay |
| Streaming request bodies | Body::into_data_stream |
onStreamingBody, with a limit per route and flow control back to the client |
| Resumable uploads | None built in; tus through other crates | GarudaUploads: the IETF resumable upload protocol |
| Interim responses | None: hyper sends only 100 Continue | response.sendInterim, such as 103 Early Hints |
| WebSockets | WebSocketUpgrade |
app.webSocket, whole messages, pings and permessage-deflate by the engine (over HTTP/1.1, HTTP/2 and HTTP/3) |
| WebTransport | None in hyper | app.webTransport |
| HTTP client | reqwest | request.client, HTTP/1.1 and HTTP/2, redirects by policy, decompression, streamed responses and server-sent events, a whole-exchange budget |
| PostgreSQL | sqlx, tokio-postgres | Native driver on the poller |
| Redis | redis-rs, fred | Native driver on the poller: RESP3 and RESP2, TLS, ACL, pipelines, transactions, pub/sub |
| SQLite | sqlx, rusqlite | The system's libsqlite3 on the blocking pool: a writer and readers per worker, WAL, migrations |
| Blocking work | spawn_blocking |
blocking { … } on a bounded pool of threads per worker |
| Testing | tower::ServiceExt::oneshot |
app.test, the real engine |
Every row is implemented and tested.
For performance, BENCHMARKS.md has the method and every run, including hello-world comparisons with axum that measure what the server adds to a request.
A Tower layer awaits the handler and gets its response back as a value. A
Garuda handler writes its answer straight into the connection's buffer, so
there is no response value to hand back without building and copying one on
every request. Middleware runs before the handler, and response.onSend
lets it see and change the response just before the head is written. The hook
runs for every way a request is answered: the handler, a thrown error, a
middleware's refusal and a deadline's 504. Hooks run last-added first, so they
nest the way layers do.
What that rules out: retrying the handler from middleware, or holding a scope
open around its run. Retry inside the handler or around pool.transaction, and
use app.deadline for timeouts. Hooks do not run for static files, cache hits,
or the 404 before any route matches.
An axum layer applies only to routes added before it. Garuda's use covers
every route in its group or application wherever the call is, so moving a line
cannot leave a route unguarded. A route that must skip a middleware goes outside
the group.
A worker is a process with one thread. Nothing is shared, so nothing is locked,
a crash takes only that worker's connections, and SIGHUP replaces workers one
at a time. The cost: in-memory state is per worker, and a pool of 8 database
connections is 8 per worker. Keep shared state in a database, and size pools as
the total divided by --workers.
A crash staying in one worker matters more in Swift than in Rust: a
force-unwrapped nil or an index out of range ends a Swift process on the
spot, where Tokio catches a Rust panic in the task that raised it.
Processes cannot steal work from each other the way Tokio's threads do.
Instead a worker ahead of the others leaves new connections to them, and a
worker where quick requests would wait behind slow ones hands idle HTTP/1
keep-alive connections to one where they would not (--balance). When every
worker holds a slow client, the slow ones are gathered onto fewer workers to
free the rest. With 8 slow connections among 64 on 8 workers, starting a
second in, the quick requests' p99 was 2.4–2.8 ms and axum's 3.0–3.1 ms, where
the kernel's hash alone let it reach 7.6 ms. HTTPS connections do not move:
the record layer is BoringSSL and encrypts in the process, so the session
cannot travel with the descriptor. HTTP/2 connections, WebSockets and requests
in progress stay on the worker that has them.
A worker is one thread, and staying on it is what removes the scheduling hop. A
loop that never awaits holds the worker until it ends. A deadline bounds
waiting, not computing. Run more workers than busy cores, and hand a call that
blocks or computes for long to try await blocking { … }, which runs it on the
worker's blocking pool while the worker serves other requests. Waiting on an
external program is such a call.
Swift 6.1's concurrency runtime, the one macOS 15 has, sends a task that calls
Task.yield() to the global executor whatever executor it prefers. The
handler still resumes on its worker, but only once a thread of the global
pool is free to hand it back. Linux, and the runtime in newer macOS, keep it
on the worker. A handler that wants to let others run is better served by an
await on the engine, such as response.sleep(milliseconds:).
Garuda brings its own JSON coder, UUID and Timestamp. If your code imports
Foundation too, write Garuda.UUID where both are in scope.
Every connection keeps up to 256 statements prepared, with no opt-in per query.
Behind PgBouncer in transaction pooling mode, set statementCacheCapacity = 0.
These work without any handler code, set by flags:
- HTTP/1.1 with a parser strict where it prevents request smuggling,
HTTP/2 over TLS and cleartext, and HTTP/3 over QUIC (
--http3). QUIC, TLS 1.3 key schedule and QPACK are Swift, over OpenSSL's crypto primitives. - TLS on a vendored BoringSSL, with several certificates chosen by SNI on
both TCP and QUIC. OpenSSL is still linked for cryptography, ACME and QUIC.
There is no kernel TLS, so
--ktlsis accepted and ignored. - ACME certificates (
--acme-domain), obtained and renewed with tls-alpn-01 on the port already served. - Static files (
--static-dir) withsendfile,ETagand byte ranges, pre-compressed copies (--compress-static), and directories answered by theirindex.html(--static-index) or listed (--static-listing). - Rate limiting (
--rate-limit) counted across workers. - HTTPS redirects (
--redirect-http), HSTS, request IDs, W3C trace context, an access log, Prometheus metrics and a health check. - Graceful shutdown with
--drain-delay, and zero-downtime reload onSIGHUPor when the executable is rebuilt (--reload). - Unix sockets, multiple workers, and trusted proxy headers
(
--forwarded-allow-ips). - Compression of handler responses (
--compress) and a response cache shared by the workers (--cache-size).
garuda --help lists every flag, and CONFIG.md explains them.
| Signal | Effect |
|---|---|
SIGTERM |
Stop accepting, finish in-flight requests within --graceful-timeout. With --drain-delay, keep serving first while the health check answers 503. |
SIGINT, SIGQUIT |
Shut down without the drain delay. |
SIGHUP |
Replace every worker one at a time, rereading certificates, without refusing a connection. |
swift test # 1130 unit tests, and the fuzz corpus
(cd Examples && swift test) # 41 the examples, through app.test
bash scripts/compile-fail-test.sh # 18 handler code that must not compileThe end-to-end suites run against a release build. Each takes a binary path as
its first argument. Most use .build/release/garuda; handler-test.py,
websocket-test.py, webtransport-test.py, upload-test.py and broadcast-test.py use
.build/release/garuda-conformance, whose routes exist only for the tests.
bash scripts/integration-test.sh # 36 HTTP/1.1 framing and smuggling defences
bash scripts/static-test.sh # 105 --static-dir
bash scripts/compress-test.sh # 76 --compress, --compress-static
bash scripts/cache-test.sh # 85 --cache-size
bash scripts/ratelimit-test.sh # 18 --rate-limit
bash scripts/redirect-test.sh # 22 --redirect-http, --hsts
bash scripts/sni-test.sh # 11 certificates by SNI
bash scripts/resumption-test.sh # 7 session tickets resume, TLS 1.3 and 1.2
bash scripts/acme-test.sh # 12 --acme-domain, needs Pebble
bash scripts/request-id-test.sh # 12 --request-id
bash scripts/trace-context-test.sh # 17 --trace-context
bash scripts/drain-test.sh # 14 --drain-delay
bash scripts/reload-test.sh # 7 SIGHUP under load
python3 scripts/feature-test.py # 62 shutdown, supervision, unix sockets, slow clients
python3 scripts/http2-test.py # 62 against the h2 library
python3 scripts/http3-test.py # 76 against aioquic
python3 scripts/router-streams-test.py # 41 routes over HTTP/2 and HTTP/3
python3 scripts/handler-test.py # 143 the handler API over all three protocols
python3 scripts/websocket-test.py # 104 handshake, framing violations, closing, pings, deflate
python3 scripts/websocket-streams-test.py # 94 WebSocket over HTTP/2 and HTTP/3
python3 scripts/webtransport-test.py # 46 sessions, streams, datagrams
python3 scripts/upload-test.py # 35 streamed request bodies, 1xx, resumable uploads
python3 scripts/broadcast-test.py # 36 topics across workers, Last-Event-ID, keep-aliveThe shell suites need curl and openssl. The Python suites are clients only,
using h2 and aioquic, which share no code with the server. Two unit tests
want a second loopback address, which Linux has and macOS has to be asked for
(sudo ifconfig lo0 alias 127.0.0.2 up); without it they say so and skip,
rather than failing as though something were wrong. Every parser that
reads network bytes is fuzzed with swift run -c release pgfuzz
(fuzz/README.md). CI (.github/workflows/ci.yml) runs on
every push and pull request: the build and unit tests on Ubuntu 24.04 and
macOS 15, the handler code that must not compile, the connectors against a
real PostgreSQL and Redis, and the end-to-end suites. The protocol suites and
a sanitizer fuzz run on every push to main, nightly and on demand, but not
on a pull request: they are slow and want a QUIC stack.
Garuda is past 1.0, so a breaking change to the public API needs a major
version. 1.1.0 made one exception, one word long: CompletedUpload.digest()
must now be awaited, and RELEASE.md says why.
COMPATIBILITY.md says what that covers and how much notice
a change gets. Pin with from: "1.1.0".
A handler waiting on something other than the engine -- its own continuation,
say -- is not unwound when its request is cancelled. It resumes to find
response.isCancelled set, and anything it sends is dropped. Wrap the wait in
response.cancellable { … } and it is given up on when the request ends. What
that cannot do is stop work which ignores cancellation; the worker counts such
work and reports it on the health check instead.
- Reads from Redis replicas: every command goes to the master, or to the node that owns the slot. CONNECTORS.md has the rest of the driver's limits.
multipart/byteranges: a request for several ranges at once is answered with the whole file.- QUIC session resumption and 0-RTT.
- TLS over TCP in Swift: the record layer is BoringSSL.
- Kernel TLS: BoringSSL has none, so
--ktlsis accepted and ignored. - Windows, except through WSL 2.
| File | What it covers |
|---|---|
| HANDLER-API.md | The handler API's design, decisions and roadmap |
| INSTALLATION.md | Building, dependencies, certificates, deployment |
| CONFIG.md | Every command-line flag |
| MIDDLEWARE.md | How middleware runs, every piece Garuda ships, and writing your own |
| EXAMPLES.md | The runnable applications, and recipes: the per-worker model, settings, common tasks |
| ARCHITECTURE.md | How the engine is built |
| TRANSPORT.md | What each protocol implementation does |
| Examples/README.md | Eight runnable applications and how they are laid out |
| Examples/STARTER.md | The starter application: layout, configuration, migrations, deployment |
| CONNECTORS.md | The HTTP client and database drivers: limits and future work |
| BENCHMARKS.md | Benchmark method and results |
| COMPATIBILITY.md | What an application may depend on, and what a release may change |
| DEPLOYMENT.md | Publishing: the Swift Package Index listing, and the release checklist |
| RELEASE.md | Changes |

