uak35 / laravel-weighted-dbmanager
Smooth weighted round-robin read-replica routing, windowed writer fallback and replica health tracking for Laravel
Package info
github.com/UAK-35/laravel-weighted-dbmanager
pkg:composer/uak35/laravel-weighted-dbmanager
Requires
- php: ^8.4
- illuminate/console: ^12.0
- illuminate/database: ^12.0
- illuminate/http: ^12.0
- illuminate/redis: ^12.0
- illuminate/routing: ^12.0
- illuminate/support: ^12.0
- symfony/process: ^7.2
Requires (Dev)
- laravel/pint: ^1.25
- orchestra/testbench: ^10.0
- phpstan/phpstan: ^2.0
- phpunit/phpunit: ^11.5
Suggests
- ext-redis: Required by RedisAtomicStateStore when using the phpredis client
- predis/predis: Alternative pure-PHP Redis client for RedisAtomicStateStore
Provides
None
Conflicts
None
Replaces
None
This package is auto-updated.
Last update: 2026-09-27 17:18:01 UTC
README
Smooth Weighted Round-Robin (SWRR) read-replica routing for Laravel 12.
Laravel picks a read replica with Arr::random(), which ignores how much
hardware each replica actually has. This package swaps that pick for a
WeightedDatabaseManager that:
- resolves an integer weight per replica from
cpu_cores/ram_gb(linear or diminishing formula), or from an explicitweightkey; - runs one SWRR step per read connection — the same algorithm nginx uses for upstreams, so distribution is exact over Σweights draws and never clumps;
- keeps the SWRR state in Redis behind a Lua
EVAL, so N PHP-FPM workers share one rotation instead of drifting apart; - degrades to an in-process store when Redis is unavailable, and says so;
- tracks replica health with exponential cool-down (30 s → 30 min) and filters failing replicas out of the pool;
- can hand reads to the writer outside configurable
reader_windows, and keeppgcat.tomlin lock-step with that decision.
How the wiring works
Laravel chooses the read host in
Illuminate\Database\Connectors\ConnectionFactory::getReadConfig() — not in
DatabaseManager. So the provider does two things:
- binds
db.factorytoWeightedConnectionFactory, and - installs a resolver on it that calls
WeightedDatabaseManager::readConfigFor()for every read PDO.
readConfigFor() then decides, in order: the writer (when a TimeWindowResolver
says we are outside every reader window) → Laravel's own random pick (when the
replicas carry no cpu_cores / ram_gb / weight metadata, so upgrading is a
no-op) → SWRR over the weighted, health-filtered pool.
Binding the factory is what makes any of this run; the manager alone cannot intercept the choice.
Requirements
- PHP 8.4+ — the floor
composer.jsonenforces with"php": "^8.4" - Laravel 12 (
illuminate/console,illuminate/database,illuminate/http,illuminate/redis,illuminate/routing,illuminate/support) - Redis (
ext-redisorpredis/predis) ifswrr.primary_storeisredis— the default.localworks without Redis but only coordinates within a single process. - A host app whose replica entries carry
cpu_cores,ram_gborweight. - pgcat — optional, and PostgreSQL-only: the flipper never acts on a MySQL, MariaDB, SQLite or SQL Server connection (see Pgcat).
Installation
composer require uak35/laravel-weighted-dbmanager
Two steps follow, and neither is optional:
-
Publish the config, then edit it. The application owns this file:
php artisan vendor:publish --tag=db-manager-config
See The application owns its config for what that means in practice.
-
Register the provider. It replaces the framework's
DatabaseServiceProviderand is not auto-discovered — see Register the provider.
The application owns its config
Every install keeps its own config/db-manager.php, and that copy is the
configuration the application runs on. The file shipped inside the package is the sample it is published from.
The sample holds sample values — values that illustrate the shape of the config, not values chosen for your installation. Publishing it is therefore the start of the config step rather than the end of it: the installer is expected to open the published copy and change it. That is the intended workflow, not a workaround.
If the published copy is missing the package still merges its sample as the base, so nothing breaks — but the app is then running on sample values, which is an unfinished installation rather than a configured one.
Publishing on Composer's install and update events
Publishing is a Composer step, but Composer will not take it on the package's behalf: only the root package's scripts run, so the scripts declared in this package fire when this checkout is the root and never when an application requires it as a dependency. The application declares the call itself, beside its other publish scripts:
"post-install-cmd": [ "@php artisan vendor:publish --tag=db-manager-config --ansi" ], "post-update-cmd": [ "@php artisan vendor:publish --tag=db-manager-config --ansi" ]
Note the missing --force: vendor:publish copies a file only when the
destination does not exist and reports SKIPPED when it does. That is precisely
the behaviour this setup wants — every install ends with a config the application
owns, and no install ever writes over the copy the installer has edited.
An unknown tag is not an error either: the framework prints
No publishable resources for tag [...] and exits 0, so the line is harmless in
an application that has the package installed but not yet wired up. Deploys that
run composer install --no-scripts skip it; run the artisan command from the
release step instead.
The package ships the same call as a script of its own — the one its
post-install-cmd and post-update-cmd events run — and you can point it at an
application by hand:
composer publish-config # from a checkout of this package php bin/publish-config.php path/to/app # or at an application, from anywhere
Given a directory that is not an application, it says so and exits 0, so a stray
call is never an error. Given an application it runs vendor:publish for the tag,
and fails with a hint when the tag does not resolve — which is what happens while
the provider is still unregistered, since the tag comes from the provider's
boot().
Linking into a local application (read the package from disk)
To point an application at a working copy of this package instead of a tagged
release, register the checkout as a Composer path repository and require it at
its dev version. bin\composer-link.cmd does both, from the application root:
rem run this from the root of the Laravel application, not from the package
bin\composer-link.cmd ..\path\to\laravel-weighted-dbmanager
It reads name out of the package's composer.json, writes a
repositories.local path entry into the application's composer.json, and
requires <name>:@dev. Composer then junctions
vendor\uak35\laravel-weighted-dbmanager onto the checkout, so an edit under
src/ is live in the application with no commit and no composer update.
Anything after the directory is forwarded to composer require, so flags pass
straight through:
bin\composer-link.cmd ..\path\to\laravel-weighted-dbmanager --dry-run
--help prints the full description. COMPOSER_LINK_COMPOSER,
COMPOSER_LINK_PHP and COMPOSER_LINK_JQ override the programs it detects.
PHP reads the JSON, so unlike the bash original this needs no jq.
Undo with composer remove uak35/laravel-weighted-dbmanager and
composer config --unlink repositories.local.
Register the provider (explicit — it is not auto-discovered)
The package provider extends Illuminate\Database\DatabaseServiceProvider and
must replace it: two providers binding db means whichever registers last
wins. That is why extra.laravel.providers is intentionally empty —
auto-discovery can only add the provider, never remove the framework's.
Illuminate\Foundation\Bootstrap\RegisterProviders builds the provider list
like this:
config('app.providers') ?? ServiceProvider::defaultProviders()->toArray() // … and only then appends whatever bootstrap/providers.php returns
So bootstrap/providers.php can only append. The framework's
DatabaseServiceProvider is one of the 23 entries in defaultProviders() and
cannot be dropped from that file. Declare the swap in config/app.php instead,
built from the defaults so no other framework provider is lost:
'providers' => \Illuminate\Support\ServiceProvider::defaultProviders()->replace([ \Illuminate\Database\DatabaseServiceProvider::class => \Uak35\WeightedDbManager\Providers\WeightedDatabaseServiceProvider::class, ])->toArray(),
replace() swaps the class in place and leaves the other 22 defaults alone. It
is keyed by class name, so the key is simply ignored if the framework ever stops
listing DatabaseServiceProvider by default.
The package ships that file as a starting point: config/app.php is the Laravel
12 skeleton with exactly this block appended. It points the map at
App\Providers\WeightedDatabaseServiceProvider — an app-local provider that
extends the package one — so point the map at
Uak35\WeightedDbManager\Providers\WeightedDatabaseServiceProvider itself if you
have no app-local subclass.
Two things worth knowing:
- Registering the package provider in
bootstrap/providers.php"works" only by accident of ordering: the framework's provider registers first, so the package provider overwrites thedbbinding — leaving two providers claimingdb. Theproviderskey removes the ambiguity instead of relying on it. - With a cached config (
php artisan config:cache)RegisterProvidersskips the merge step entirely and readsapp.providersfrombootstrap/cache/config.php. The swap still applies; re-runconfig:cacheafter editing the key.
Configuration
The application owns config/db-manager.php. Every install keeps its own copy and
edits that; the package's copy is the sample it is published from.
| File | Role |
|---|---|
<app>/config/db-manager.php |
the configuration — installed once, edited per app |
vendor/uak35/laravel-weighted-dbmanager/config/db-manager.php |
the sample vendor:publish copies, and the fallback when the app's copy is missing |
The two are combined by Laravel's mergeConfigFrom(), which merges shallowly with
the app's copy winning. Because the whole config is a single nested swrr block,
an app copy shadows the sample completely rather than key by key: what your
config/db-manager.php contains is what the app runs on.
The swrr block is global; per-connection tunables stay on the connection itself,
next to the read list being weighted.
| Key | Env | Sample value | Meaning |
|---|---|---|---|
swrr.primary_store |
DB_STORE_PRIMARY |
redis |
redis or local |
swrr.redis_connection |
SWRR_REDIS_CONNECTION |
default |
Connection name used by RedisAtomicStateStore |
swrr.state_ttl |
SWRR_STATE_TTL |
86400 |
TTL (s) for the SWRR state key |
swrr.key_prefix |
SWRR_KEY_PREFIX |
swrr |
Key prefix, shared by both stores |
swrr.allow_local_fallback |
SWRR_ALLOW_LOCAL_FALLBACK |
true |
Fall back to the in-process store instead of throwing. Read as a switch — on/off, 1/0, 'on'/'off', 'yes'/'no' — and a value that is none of those is refused rather than cast |
swrr.default_weight_cpu_factor |
DB_DEFAULT_WEIGHT_CPU_FACTOR |
3.0 |
CPU multiplier when a connection sets none |
swrr.default_weight_ram_factor |
DB_DEFAULT_WEIGHT_RAM_FACTOR |
3.375 |
RAM multiplier when a connection sets none |
swrr.default_weight_formula |
DB_WEIGHT_FORMULA |
linear |
linear or diminishing |
swrr.reader_windows |
— | 10:00–14:20, 17:00–20:30 |
List of ['start' => 'HH:MM[:SS]', 'end' => 'HH:MM[:SS]'] where readers are used — an entry that is not that shape is refused, never read as a window |
swrr.reader_days |
— | [1,2,3,4,5] |
ISO-8601 days readers are allowed on; one day number, or an array of them ('1,2,3' as a single string is refused) |
swrr.timezone |
— | UTC |
Timezone for window evaluation |
swrr.pgcat |
SWRR_PGCAT_* |
(see Pgcat) | Pgcat flipper settings — PostgreSQL connections only |
The sample sets two weekday windows — 10:00–14:20 and 17:00–20:30, read
in swrr.timezone (UTC) — so reads are served by the weighted replica pool
inside them and by the writer for the rest of the day. Leave reader_windows
empty or unset and TimeWindowResolver is permissive instead: reads always go to
the replica pool. Either way reads are served by the writer on non-reader days
and outside every window (start inclusive, end exclusive), and a window like
['start' => '00:00:00', 'end' => '00:00:00'] can never match, which is a
convenient way to force writer-only mode.
A window written any other way is refused, not repaired. An entry that is not an
array — the flat '10:00-14:20' is the first one everybody writes — is left out of the
windows the resolver is built with, logged at error level by every boot that sees it,
and failed by db:doctor's reader windows row, so --strict rejects it in a
pipeline. Reading it silently as no window at all is what used to happen, and it is
the worst of the options: it inverts the setting, so reads use the replica pool at every
hour while the config says the writer covers the margins. Well-formed entries beside a
refused one still apply, and the message says how many, so a typo costs the entry rather
than the fallback. Nothing throws, so an installation with a typo boots and serves — it
just stops doing so quietly. The decision and the alternatives are recorded in
docs/reader-windows-refusal.md.
swrr.reader_days is refused the same way, for the same reason. A day list is a list of
day numbers — [1, 2, 3, 4, 5], or the single '3' a scalar .env value produces — and
'1,2,3', which is how a list is written in .env, is not one. It used to be dropped
entry by entry, leaving no day at all, and no day leaves the resolver permissive: a day
list meant to restrict reads runs as its opposite, all week. It is logged at error
level by every boot that sees it under swrr.reader_days.refused, and db:doctor's
reader windows row FAILs on it, naming the value and quoting the accepted shape. A
list with one bad entry keeps the entries that are days, and says how many.
A boot reports every reader finding it can, not the first one it finds: a flat window
string beside a comma-separated day list is two log lines and two records, each closed out
by the boot that sees its own fix. One key is one finding, though — the record keeps one
entry per key — so the two ways the pool ends up unused (no day in 1…7, no window that can
ever be entered) are one sentence when both hold. That rule is enforced rather than assumed:
BootAudit logs every finding a boot produced before it remembers any of them, and two
findings claiming one key are logged for what they are — a defect in the package — with the
louder of the two kept for the record, so a branch that reuses a key cannot cost another
finding its log line. Nothing is claimed about windows when the day list leaves the resolver
permissive: mode is decided before a window is read, so those reads cannot be wrong about
one. The rule and the alternatives are in
docs/boot-audit-finding-keys.md.
Switches are read, not cast. swrr.allow_local_fallback, swrr.pgcat.enabled and
swrr.pgcat.use_reload are on/off settings, and PHP's own rule for turning a value into
a boolean makes 'false', 'off' and 'no' — the three ways everyone writes off —
into true. On swrr.pgcat.enabled that is not a cosmetic slip: it arms a file swap the
operator asked for the opposite of. So each is read from a closed list of spellings —
true/false, 1/0, 'on'/'off', 'yes'/'no', case and surrounding space ignored
— and anything else is refused: the setting resolves to the value its own documentation
prints (the flipper stays off, a flip reloads, local fallback stays on), the boot logs it at
error level under its own key (swrr.pgcat.enabled.refused, swrr.pgcat.use_reload.refused,
swrr.allow_local_fallback.refused), and db:doctor's switch values row fails on it.
That is also why the sample config does not wrap those values in (bool) env(...): the
cast upstream would swallow the typo before the package could refuse it. The decision and the
alternatives are recorded in docs/switch-values.md.
Each key can also come from .env: the sample resolves them through env(), and a
published copy carries those calls across, so .env stays the place for
per-environment values while the file stays the place for everything else.
Upgrading
Shallow merging cuts both ways. A key added to the sample by a newer version does
not reach an app whose copy predates it, because that copy's swrr block replaces
the sample's wholesale. After a composer update, compare the two and carry any new
keys across:
# Unix diff -u config/db-manager.php vendor/uak35/laravel-weighted-dbmanager/config/db-manager.php # Windows fc config\db-manager.php vendor\uak35\laravel-weighted-dbmanager\config\db-manager.php
Never reach for vendor:publish --force on a config you have edited: it overwrites
your copy with the sample.
Weighting a connection
'pgsql' => [ 'driver' => 'pgsql', 'write' => [[ 'host' => '127.0.0.1', 'port' => 5433, /* … */ ]], 'read' => [ ['host' => '10.0.1.11', 'port' => 5432, 'cpu_cores' => 16, 'ram_gb' => 64], ['host' => '10.0.1.12', 'port' => 5432, 'cpu_cores' => 8, 'ram_gb' => 32], ['host' => '10.0.1.13', 'port' => 5432, 'weight' => 0], // explicitly disabled ], // optional per-connection overrides 'weight_formula' => 'diminishing', 'weight_cpu_factor' => 3.0, 'weight_ram_factor' => 3.375, ],
linear → round(cores × cpu_factor + ram_gb × ram_factor).
diminishing → round(cores^0.7 × cpu_factor + √ram_gb × ram_factor), for clusters
spanning very different hardware tiers.
Pgcat
pgcat is a PostgreSQL connection pooler. It
sits between the app and the database and decides which server a query reaches, so
when reader_windows moves reads between the replica pool and the writer, pgcat has
to be told as well. PgcatConfigFlipper does that — it copies the variant of
pgcat.toml matching the current mode into place, signals pgcat, and records the
mode it applied. db:pgcat-flip drives it.
It only ever acts on a PostgreSQL connection
pgcat proxies PostgreSQL and nothing else. The provider therefore tells the flipper
which connection is current (database.default) and which driver it uses, and the
flipper refuses to act on a driver it cannot front — MySQL, MariaDB, SQLite, SQL
Server. There is no pool to swap and no pgcat.toml to keep in step, so on those
connections:
| Surface | Behaviour |
|---|---|
status()['enabled'] |
false |
status()['configured_enabled'] |
what swrr.pgcat.enabled alone says, so an explicit setting is not hidden |
status()['connection'], ['driver'], ['driver_supported'] |
name the connection and why it is off |
status() |
the full snapshot: everything below plus the file paths and commands |
disabledReason() |
the sentence explaining it — null whenever pgcat does apply |
armedReason() |
the opposite sentence — names both conditions that passed, the switch and the driver; null whenever pgcat does not apply |
isMismatched() / armingWarning() |
the states that read as armed but cannot act: swrr.pgcat.enabled = true on a connection pgcat cannot front (mismatch is true), or an armed flipper missing one of its three paths (a flip would throw). The flipper stays inert either way, and the warning names what to change |
db:pgcat-flip, --watch, --force-mode |
the command prints that reason and exits 0 at once; it never enters the watch loop, so a scheduled or supervised run costs nothing |
applyCurrentState() / forceMode() |
a no change result carrying the same reason, returned before the lock, the file copy or supervisorctl |
db:pgcat-flip --status |
prints all of it, including the inactive reason, the armed reason and any warning |
db:replica-status |
the same snapshot as a Pgcat: line — active with driver, connection and mode, or inactive and why — plus a Warning: line when it is expected to act but cannot |
GET /health/db |
the same snapshot in a pgcat block (see Health endpoint) |
| the application log | two [WeightedDB] lines, at boot: a warning per process while the setting and the driver disagree, and a warning on the boot that sees it corrected — the first state would otherwise never be mentioned, the second would never be closed out |
PostgreSQL is recognised as pgsql, postgres or postgresql, matched
case-insensitively. A flipper constructed without a driver — which is what the unit
tests do — keeps its unrestricted behaviour.
Being inert is not the same as being wired correctly, so the two are reported
separately: reason says why it will not act, armed_reason says why it will, and
exactly one of them is null. When the setting and the driver disagree —
swrr.pgcat.enabled = true on a connection pgcat cannot front — mismatch is
true and armingWarning() carries the sentence, because the flipper will never
flip however long it runs.
That case is also logged, once per process at boot, so it is visible without
anyone reading /health/db:
[WeightedDB] swrr.pgcat.enabled is true but pgcat will never act: connection "sqlite"
uses driver "sqlite", and pgcat only fronts PostgreSQL. Set swrr.pgcat.enabled = false,
or point database.default at the pgcat-fronted connection. {"connection":"sqlite","driver":"sqlite"}
"Once per process" is the point: Octane reuses a worker, so it is one line per
worker rather than one per request or per flip. The guard is static, so a process
that boots several applications — the test suite, config:cache — still logs it
once. An armed flipper whose paths are not mounted yet is transient and is not
logged; a flip reports that as a failure on its own.
The other half of the story. Correcting the mismatch means editing configuration, and configuration is read once per process — so the boot that sees the fix is a different process from the one that logged the warning, and a static guard cannot tell it that there is anything to resolve. The unresolved mismatch is therefore written down, beside the flipper's own state file:
$TMPDIR/pgcat-flip-state.json the flipper's last applied mode
$TMPDIR/pgcat-flip-state-gate.json the mismatch that has not been resolved yet
{
"gate_state": "mismatch",
"connection": "sqlite",
"driver": "sqlite",
"reported_at": "2026-09-25T12:25:15+00:00"
}
The record lives in dirname(swrr.pgcat.state_file), so it follows whatever
persistence that path was given — a /var/lib/… state file survives a reboot, the
sys_get_temp_dir() default does not. It exists only while a mismatch is
unresolved: the boot that finds the gate open while it stands logs the resolution
and deletes it, and an installation that was never mismatched never creates it. The
two lines then sit next to each other in the log:
[WeightedDB] swrr.pgcat.enabled is true but pgcat will never act: connection "sqlite" uses driver "sqlite", and pgcat only fronts PostgreSQL. … {"severity":"warning","connection":"sqlite","driver":"sqlite"}
[WeightedDB] Pgcat flipping is active, so the earlier mismatch no longer applies. {"severity":"warning","resolved":true,"connection":"pgsql","driver":"pgsql","mismatch_reported_at":"2026-09-25T12:25:15+00:00"}
The resolution is a warning for the same reason the store's recovery is: window
flips start happening from that boot on, which is worth noticing. It is logged once
per transition rather than once per boot, the record is cleared so that a second
mismatch is reported as a new problem instead of being resolved by the boot after
it, and switching pgcat off on purpose resolves nothing — the record stays, so a
later, correct arming still closes the warning out. Both writes are silenced and
never fatal: a read-only deploy loses the pair of lines, not the application, and
db:doctor's pgcat files row reports that directory as unwritable so the loss is
visible where someone can act on it.
Flipper settings
The pgcat block lives inside swrr, and every key is optional. The sample ships
it switched off — flipping is a deliberate step, so turn it on in your own copy
once the paths below point at your pgcat install:
'pgcat' => [ // No (bool) here on purpose: a cast reads 'off' as on, so `enabled` and // `allow_local_fallback` hand over the value as written and let the package // read it — see "Switches are read, not cast" above. 'enabled' => env('SWRR_PGCAT_ENABLED', false), 'config_path' => env('SWRR_PGCAT_CONFIG', '/etc/pgcat/pgcat.toml'), 'readers_path' => env('SWRR_PGCAT_READERS', '/etc/pgcat/pgcat-readers.toml'), 'no_readers_path' => env('SWRR_PGCAT_NO_READERS', '/etc/pgcat/pgcat-no-readers.toml'), 'restart_command' => 'supervisorctl restart "pgcat:*"', 'reload_command' => 'supervisorctl signal HUP "pgcat:*"', 'use_reload' => true, 'state_file' => sys_get_temp_dir().'/pgcat-flip-state.json', 'lock_file' => sys_get_temp_dir().'/pgcat-flip.lock', ],
| Key | Env | Sample value | Meaning |
|---|---|---|---|
pgcat.enabled |
SWRR_PGCAT_ENABLED |
false |
Master switch, still subject to the PostgreSQL-only rule above. The sample leaves it off |
pgcat.config_path |
SWRR_PGCAT_CONFIG |
/etc/pgcat/pgcat.toml |
The file pgcat reads |
pgcat.readers_path |
SWRR_PGCAT_READERS |
— | Variant applied while readers are in use |
pgcat.no_readers_path |
SWRR_PGCAT_NO_READERS |
— | Variant applied in writer-only mode |
pgcat.restart_command |
— | supervisorctl restart "pgcat:*" |
Run after the swap — the program name has to match what supervisord runs |
pgcat.reload_command |
— | supervisorctl signal HUP "pgcat:*" |
Used instead when use_reload is true |
pgcat.use_reload |
— | true |
Prefer a HUP over a full restart. Read as a switch like allow_local_fallback, so 'no' is off and a value that is neither is refused |
pgcat.state_file |
— | sys_get_temp_dir() |
Last applied mode, so a poll that changes nothing copies nothing — the boot check keeps its unresolved-mismatch record beside it |
pgcat.lock_file |
— | sys_get_temp_dir() |
Mutex — a concurrent flipper returns skipped |
The swap is atomic: the variant is written to {config_path}.tmp.{pid} and then
rename()d, so pgcat never reads a half-written file; only then is the restart or
reload signal sent. Either transition is one copy plus one signal.
A flip also judges its supervisor command before it replaces anything, and puts the
previous file back if the command fails anyway. Both exist for the same state, which is
the one an operator cannot see: a new pgcat.toml on disk while the running pgcat
processes still hold the old one. Nothing is wrong yet — and then pgcat restarts for some
unrelated reason, and comes up in a mode no window asked for.
| When the command cannot work | What a flip does |
|---|---|
| before the swap | refuses: nothing is copied, nothing is run, no mode is recorded — so the next poll retries and the flip goes through by itself once the command is fixed |
| after the swap (it passed the check and failed anyway) | rolls the previous variant back, so disk and running processes agree again |
The check is the same read-only question db:doctor asks — supervisorctl status "pgcat:*", never a restart or a signal — so the row and the flip cannot disagree about
whether a command works. It costs one process per flip attempt, and only when a flip is
actually about to happen. A rollback that itself fails says so, loudly, in the flip's
error: that is the one outcome that needs a human.
That last step is a shell command line, so two things about the name in it matter.
It has to match what supervisord runs — "pgcat:*" addresses every program in a group
called pgcat (which is what a [program:pgcat] section becomes), and a single
program can be named on its own — and it has to be quoted, because an unquoted
pgcat:* is a glob: the shell may rewrite it before supervisorctl ever sees it, and
what arrives is then decided by the working directory at the moment of the flip.
php artisan db:doctor's pgcat supervisor row judges both of those, plus whether
supervisorctl resolves for the user the flipper runs as and whether supervisord knows
that name — see Preflight.
Rehearsing a flip: --dry-run
A window flip is a file replaced and a process restarted, and the person who wants to
know whether that will work is rarely the person who is allowed to do it. --dry-run
performs every step of a flip that leaves nothing behind and reports the two that do:
$ php artisan db:pgcat-flip --dry-run
[dry run] would flip: never → readers (nothing was renamed, and pgcat was neither signalled nor restarted)
lock done taken and released: /var/lib/lpr/pgcat-flip.lock
read source done /etc/pgcat/pgcat-readers.toml (1892 bytes)
supervisor check done supervisorctl resolves to /usr/bin/supervisorctl and knows "pgcat:*": pgcat:pgcat_00 RUNNING pid 4242, uptime 0:12:34
write temp done /etc/pgcat/pgcat.toml.tmp.4711 (1892 bytes)
remove temp done the rehearsal leaves nothing behind
rename would /etc/pgcat/pgcat.toml.tmp.4711 → /etc/pgcat/pgcat.toml (the content would change)
supervisor would supervisorctl signal HUP "pgcat:*"
state file would /var/lib/lpr/pgcat-flip-state.json (writable, left as it is)
The source is read for real, the flip's own supervisor check is run for real (it is
read-only — the same supervisorctl status "pgcat:*" the flip makes before it replaces
anything), and the temp file the swap writes is written for real into the target's own
directory — then removed again. That write is the one thing a metadata check cannot do:
db:doctor's pgcat files row judges the directory with is_writable and deliberately
leaves no trace, and a directory can pass that and still refuse the write. Proving the
write also proves the rename's precondition, since the two are on the same filesystem;
the rename itself is reportable and cannot be rehearsed without doing it.
Three things it does not do, and says so:
| Not taken | Why |
|---|---|
| the rename | the point of a rehearsal |
| the supervisor command | nothing is restarted, signalled or stopped — the command a flip would run is printed instead |
| the state write | recording a mode that was never applied would make the next real flip believe it had already happened |
The state file's writability is still answered, from metadata, because a flip's write to it is silenced: an unwritable one produces a flip that reports success and then repeats on every poll, restarting pgcat each time. The rehearsal calls that a failure even though the flip would technically succeed — that is the outcome worth catching before the window.
Exit codes are the question, not the outcome: 0 when a flip would happen or when
there is nothing to do (mode unchanged, flipper not armed, lock held — a flip would skip
those too), and 1 only when a step a flip needs did not work. That makes it usable as a
preflight on a deploy that is not allowed to flip anything yet:
php artisan db:pgcat-flip --dry-run || exit 1 # would the next flip work?
--dry-run --force-mode=writer rehearses the forced flip exactly as --force-mode
applies it, --status wins when both are passed (a state report is what was asked for),
and --dry-run --watch is refused rather than half-honoured: a watch daemon whose every
pass is a rehearsal looks exactly like one that is flipping.
--json reports any of it as one object instead of the rendered report — the verdict, the
exit code and the evidence, for a pipeline to assert on rather than to grep: see
the JSON report.
The flip itself follows the same mapping, which is written out as a table of its own
under Flipping below: applied, nothing to do and skipped all
exit 0, and only a step that did not work — or a source config that is not there — exits
1. A refused flag combination exits 1 without reaching the flipper at all.
One trace is left on purpose: the lock file, because a flip takes that lock too and cannot take it if the directory will not accept it. It is the same file every flip leaves behind.
Which steps can be rehearsed and which cannot, what each failure means and why a rehearsal
is a command rather than a db:doctor row is written up in
docs/pgcat-flip-dry-run.md.
Degradation, and why it is not silent
When the primary store is unhealthy, the manager serves from an in-process
LocalStateStore and logs an error once per process:
[WeightedDB] Degraded to the in-process SWRR store. Read rotation is no longer
shared across PHP workers until the primary store recovers.
When the primary store answers again it logs a warning and releases the
fallback. healthSummary() — and therefore /health/db and
db:replica-status — reports which store is actually serving:
| field | meaning |
|---|---|
store |
the store serving requests right now |
store_healthy |
health of the primary store |
primary_store |
the configured primary store (redis(default)) |
degraded |
true once this process has fallen back |
A worker serving from the in-process store is not sharing a rotation with its
siblings, so it is worth alerting on. Set
swrr.allow_local_fallback = false to hard-fail instead of degrading:
$manager->setAllowLocalFallback(false);
A store that is not there yet is not a failing store
Connection configs are built while providers register — any application that touches
the schema in register() does exactly that — so the first replica choice of a
process can happen before the application has bound redis. That pick falls back
like any other failure, but it is deliberately not counted as one: three early picks
would otherwise mark the store unhealthy, and that worker would stop using Redis for
the rest of its life. One warning records it per process:
[WeightedDB] This pick fell back to the in-process store.
{"store":"redis(default)","connection":"pgsql","reason":"The redis(default) store
cannot be used: the container has no \"redis\" binding yet."}
The timing of that probe — how often an installation may PING its store, and why a
static guard cannot bound it — is recorded in
docs/boot-audit-store-probe.md, together with the
mechanisms that were rejected and the measurements behind the choice.
The store reaches Redis through the container rather than the Redis facade, and
that is not a style preference. Resolving an unbound redis does not fail: the
phpredis extension ships a class called Redis and PHP class names are
case-insensitive, so the container builds the extension class — and the Redis
facade caches whatever it is handed as its root for the rest of the process. A
correct installation then throws
Error: Call to undefined method Redis::connection()
from any code in that worker, including Horizon, queues and the cache. If you see
that message, something in the application — not this package — used Redis:: before
the Redis provider registered.
Health endpoint
The package ships DatabaseHealthController, which reports per-connection replica
weights, share %, health, the active formula, store state and whether pgcat
flipping is active. Route it yourself and protect it:
use Uak35\WeightedDbManager\Http\Controllers\DatabaseHealthController; Route::get('/health/db', [DatabaseHealthController::class, 'index']) ->middleware('auth.basic');
Alongside the per-connection replicas map it carries one pgcat block — the same
snapshot db:replica-status prints and db:pgcat-flip --status shows, so the three
cannot disagree about whether flipping is active:
{
"status": "ok",
"replicas": {
"pgsql": {
".": "."
}
},
"pgcat": {
"enabled": true,
"configured_enabled": true,
"connection": "pgsql",
"driver": "pgsql",
"driver_supported": true,
"resolver_mode": "readers",
"last_mode": "readers",
"reason": null,
"armed_reason": "pgcat flipping armed: swrr.pgcat.enabled = true and connection \"pgsql\" uses driver \"pgsql\"",
"mismatch": false,
"warning": null
},
"audit": {
"available": true,
"count": 1,
"severity": "error",
"counts": {"error": 1, "warning": 0},
"oldest": "2026-09-21T08:15:00+00:00",
"findings": [
{
"key": "swrr.reader_windows.refused",
"level": "error",
"warning": "reader_windows[0] is \"10:00-14:20\". Refused: each window must be an array …",
"resolution": "swrr.reader_windows is readable again: every entry is a window …",
"first_reported_at": "2026-09-21T08:15:00+00:00",
"age_seconds": 345600,
"age": "4 days",
"context": {"rejected_windows": [{"at": "[0]", "entry": "\"10:00-14:20\""}]}
}
],
"error": null
},
"errors": {}
}
The audit block is the boot audit's standing findings: the settings that read as on
but cannot act, each with the sentence the boot log carried for it, the level it was
logged at, and how long it has stood — age in words, with age_seconds for a monitor
that would rather do the arithmetic. count and oldest save a dashboard from walking
the list.
severity and counts are the machine-readable half, and the reason the block is
alertable at all: a level per finding can only be used by a reader that walks the list and
compares strings. severity is the loudest level standing — error when at least one
finding is a value the package refused (malformed input it will not guess at, which is
what the installation is running without), warning when the loudest is a setting that
reads as on but cannot act, and none when nothing stands — so the distinction between the
two kinds of finding is one field:
| alert | rule |
|---|---|
| page — a refused value is standing | severity == "error" (or counts.error > 0) |
| ticket — a setting reads as on but cannot act | severity == "warning" |
| the audit stopped reporting | available == false |
counts carries the same distinction per level, with both keys always present so a rule
can be written as counts.error > 0 rather than as a lookup that might be missing. Levels
are compared exactly: the audit logs at two of them, and anything else in a hand-edited
record reads as the quieter one rather than being upgraded into a claim the record cannot
support.
The same rule is in the boot log, so a monitor that would rather read lines than poll every
host does not have to poll this route at all: every line the boot audit writes carries the
level it was written at as severity, in the payload beside the finding key — a refused
value still standing is severity: "error", the same comparison this route's severity
answers, and the finding is logged on every boot while it stands. It is in the payload rather
than left to the channel's own level because how a level is spelled belongs to whatever
handler is configured (production.ERROR in the default line log, level_name under a JSON
formatter, a priority number in syslog), and because ERROR there is every error the
application ever logs. The line that closes a finding is written at warning whatever the
finding stood at, and carries resolved: true — arriving at a working configuration is good
news either way — so a pager written per finding key stops on the very line that repairs it.
The decision, and the seven shapes it was weighed against, are in
docs/boot-audit-log-severity.md.
available is false when no record could be read, and error — present either way,
null whenever the record was read — says which of the two that was: null with
available false means the package's provider is not registered, and a set error means
the file could not be read. Neither is a claim about the installation, so the summary stays
severity: none.
The same list is printed by db:replica-status, so a misconfiguration is visible from
a terminal as well as from a dashboard — and both surfaces render one accessor,
BootAudit::reported(), so they cannot disagree about the record either, including when
there is none: the command prints Audit: not registered where the payload answers
available: false with error: null, instead of the silence that used to leave "no audit
installed" and "an audit with nothing to say" indistinguishable in a terminal. It
deliberately does not move this endpoint's
status: a configuration that cannot act is not the same claim as a database that
cannot answer, and whatever polls this route should keep believing the second one.
Where a standing finding is published, which surfaces were rejected (a command of its own, a second endpoint, an eviction-worthy status code) and why is recorded in docs/boot-audit-surfaces.md.
reason is the sentence to alert on when enabled is false, and null whenever
flipping is active. armed_reason is its mirror image — the proof of why it is
active, naming the switch and the driver that both passed. Exactly one of the two is
null.
The one combination that deserves attention of its own is a switch that is on for a
connection pgcat cannot front. The flipper is correctly inert — there is no pool to
swap — but nothing will ever flip, so mismatch is true and warning carries the
sentence:
{
"status": "ok",
"pgcat": {
"enabled": false,
"configured_enabled": true,
"connection": "sqlite",
"driver": "sqlite",
"driver_supported": false,
"reason": "pgcat is PostgreSQL-only; connection \"sqlite\" uses driver \"sqlite\", so pgcat flipping is disabled",
"armed_reason": null,
"mismatch": true,
"warning": "swrr.pgcat.enabled is true but pgcat will never act: connection \"sqlite\" uses driver \"sqlite\", and pgcat only fronts PostgreSQL. Set swrr.pgcat.enabled = false, or point database.default at the pgcat-fronted connection."
}
}
A flipper that cannot act does not make the endpoint unhealthy — on MySQL there
is nothing to flip, so status stays ok — and when the provider is not registered
the block is null.
Artisan commands
All four are registered by the provider.
| Command | Purpose |
|---|---|
php artisan db:doctor {connection=pgsql} [--strict] [--json] |
Preflight: check an installation end to end before traffic arrives |
php artisan db:replica-status {connection=pgsql} |
Table of host, weight, share %, health, plus formula, store backend, pgcat state and the boot audit's standing findings |
php artisan db:probe-replicas {connection=pgsql} |
SELECT 1 against every replica and feed results to HealthMonitor (schedule every 30 s) |
php artisan db:pgcat-flip [--status|--watch|--dry-run|--force-mode=] [--interval=] [--json] |
Keep pgcat.toml in step with the reader/writer window — a no-op unless the current connection is PostgreSQL |
// routes/console.php Schedule::command('db:probe-replicas')->everyThirtySeconds()->withoutOverlapping()->runInBackground(); Schedule::command('db:pgcat-flip')->everyMinute()->withoutOverlapping(60)->runInBackground();
Each probe opens a throwaway single-host connection (the pooled read/write
lists are stripped) and purges it afterwards, so probing never feeds the weighted
pool it is measuring. Diagnostics go to the replica's own host:port key, the
same key WeightResolver uses.
Probing: db:probe-replicas
The exit code is what a scheduler reports, so it says whether the sweep did its job — not whether every replica is well:
| what the sweep found | exit |
|---|---|
| at least one replica answered | 0 |
the connection has no read list at all |
0 — nothing was asked of it |
no entry in read is a replica map |
1 — nothing could be probed |
| every replica failed | 1 — nothing was marked healthy |
| the container has no weighted manager | 1 — there was no sweep |
A partial failure stays 0: marking a replica that is down is what the command
is for, and the health monitor — not the exit code — is that record. A sweep that
reached nothing is the different case, and the two ways that can happen are named
separately in the run's closing line, because they are repaired differently:
Nothing could be probed: the read list on [x] holds no replica array … is a
configuration shape (read must be a list of maps, or one map), while No replica answered: 2 probed, all failed … is reachability.
read written as a single config map rather than a list of them is one replica,
which is what routing does with it. Reading it as nothing would make the command
exit 1 every thirty seconds on an installation whose single replica is serving
fine, and an exit code that is always non-zero is an exit code nobody reads.
Every row is pinned by
DbProbeReplicasCommandTest::test_the_exit_code_is_a_function_of_what_the_sweep_reached,
which probes a real connection per row and asserts the code, the sentence beside
it, and the host:port results that had to reach the health monitor — the half of
the command that is otherwise invisible, because the probe's own connection is
purged after every attempt. That matrix is not left to be kept in step with the
table above by hand: DbProbeReplicasCommandTest::test_the_matrix_agrees_with_the_readme_exit_table
reads this table out of the README and asserts each row's number against the cell the
matrix enforces, and asserts that the two name the same cases — so a documented case
cannot lose its last cell, and a cell cannot be added without documenting it.
The flipper only ever acts on a PostgreSQL connection — Pgcat has the details.
Preflight: db:doctor
php artisan db:doctor answers the question a deploy actually cares about — will
this installation do what it is configured to do? Eleven things can be wrong while
the application still boots and answers requests, and each one is a row:
| Row | Fails when |
|---|---|
provider swap |
the framework's DatabaseServiceProvider is still in charge, so db is not the weighted manager and reads are unweighted |
weighted factory |
db.factory is not WeightedConnectionFactory — the replica is chosen there, so weighting is not running either |
published config |
warns when config/db-manager.php is missing (the package sample is in use) or still byte-identical to the sample (nobody has reviewed a value) |
pgcat gate |
swrr.pgcat.enabled is on where pgcat cannot act, or armed while a path a flip needs is unset. Reads the boot record as well as this boot's verdict, so it also fails with the age of a mismatch an earlier boot recorded, and warns when one is on record for a connection this run does not inspect. Prints swrr.pgcat.enabled = false as a suggestion for the switch armed where pgcat cannot act, and nothing for the other two failures: the paths it names are this installation's to choose |
pgcat files |
the flipper is armed and a file a flip needs is missing, unreadable or unwritable — including the target's directory, where the atomic swap writes {target}.tmp.{pid}, and the state/lock directory. The only row that never prints a suggestion: every problem here is a path or a permission, and the right value is whatever this installation's pgcat and supervisor actually use |
pgcat supervisor |
the flipper is armed and the command it runs after the swap cannot work: its executable does not resolve for this user (not an absolute path, not on PATH), the program name is an unquoted glob the shell may rewrite, supervisorctl cannot answer, or supervisor does not know the program. Read-only: the only command it runs is the flip's own with status in the verb position. Prints the setting line to paste as a suggestion for the two faults that reduce to a command — an unquoted program name, and a command left empty — and nothing for the rest, which are repaired in a PATH, a running supervisord or a [program:] section |
replica metadata |
fails when a replica's weight/cpu_cores/ram_gb is a value the resolver does not read as written — not a number, or a number under the floor that key is read at — naming the replica and what the resolver reads instead. A weight it cannot read is read as 0, and 0 is how a replica is disabled, so that replica leaves the pool: without this the row reports the smaller pool as the installation. Also names a replica disabled with weight: 0, which is a choice and does not fail. Warns when the replicas carry no metadata at all, so the resolver picks at random while the table still looks weighted |
switch values |
Fails when swrr.pgcat.enabled, swrr.pgcat.use_reload or swrr.allow_local_fallback is written as something that is not on or off — 'false', 'off' and 'no' are how off is written, and a cast reads every one of them as on, which on the pgcat switch arms the file swap. The row quotes the value, quotes the accepted spellings, and names the value the setting falls back to. Names every switch it refuses rather than the first, and dates each from its own finding in the boot record. Nothing is suggested: re-spelling 'flase' would be guessing at what was meant |
reader windows |
swrr.reader_windows holds an entry that is not a window — the flat '10:00-14:20' is the one everyone writes first — or is not a list of windows at all, or swrr.reader_days is not a list of day numbers — '1,2,3' written as one string is the same mistake one key over. Fails on both, naming the entry and quoting the shape the setting reads: reading a typo as "nothing" is what inverts the setting, so the package refuses it instead. Warns on the three ways a well-formed setting still cannot act — an empty day list (permissive), days outside 1…7 (the pool is never used), and windows whose start is not before their end (never entered). When the refused value names its own replacement — the flat string, '1,2,3' — the row also prints the setting line to paste as a suggestion under it. Names every problem it finds rather than the first, and dates each from its own finding in the boot record |
store probe |
warns when the store check can never run: probing switched off (swrr.audit.store_probe_seconds = 0), or an in-process primary store with nothing to reach. Fails when the audit record cannot be written — the record is what throttles the probe, so without it the PING is skipped on every boot and an unreachable store goes unreported. Names every state it finds rather than the first: an installation that is both switched off and unable to write the record is told both, and the row's verdict is the loudest of them |
store reachability |
the primary store cannot serve a read: no usable redis binding, or a connection that refuses — workers then fall back to an in-process rotation that is not shared between them |
Read-only, so it is safe in a deploy pipeline: nothing is written, no replica is
queried and no flip is attempted. Two checks reach outside the process: the store
check, which is a real PING on swrr.redis_connection rather than a look at the
store's own health
flag — that flag is only lowered after an operation has already failed, so on a
fresh process it says "nothing has gone wrong yet", which is not "the store is
reachable". An unreachable store therefore costs one connect timeout.
The pgcat file check is the other one that has to look past the configuration: the
gate only asks whether the three paths are set, so an installation can be armed,
gated correctly and still be unable to flip. The check reads the paths off the
flipper's own status, so it judges the files a flip really touches — the reader
source and writer source must be readable (swap() picks one by mode), the target
must exist and be writable, and the target's directory must accept the temp file the
atomic swap writes. The state_file and lock_file directories are included because
both writes are silenced with @: if they fail, a flip still reports success and
then restarts pgcat again on the next poll, forever. Every unusable file is named in
the row. Readability and writability come from is_readable/is_writable, so no
probe file is ever created.
Exit 0 when nothing failed, 1 when a check failed — or, with --strict, when any
check warned:
| the rows | --strict off |
--strict on |
|---|---|---|
| no failures and no warnings | 0 |
0 |
| warnings, no failures | 0 |
1 |
| any failure | 1 |
1 |
--strict moves the exit code and the closing line, and nothing else — the rows are the
same either way — so php artisan db:doctor --strict is a release gate. That line is what
an operator reads beside the code, so a warning the flag turned into a failure says so,
and names the flag that did it: exit 1 with nothing failed is otherwise undiagnosable
from rows that are all PASS and WARN. A pipeline reads the same three facts as fields —
exit_code, verdict and strict — from --json, a few paragraphs below.
$ php artisan db:doctor --strict
…
11 checks: 10 passed, 1 warning, 0 failed
No failures — but --strict counts the 1 warning above as failures: this run exits 1 as a release gate, not because the installation is broken.
The table above is read back out of the README by
DbDoctorTest::test_the_matrix_agrees_with_the_readme_exit_table, which asserts every cell
of the matrix against the number the table names for its case — the documented rule and the
enforced rule cannot drift apart in either direction. What the cells do is pinned by
DbDoctorTest::test_the_exit_code_is_a_function_of_the_rows_and_the_strict_flag, which
builds each row shape from real misconfiguration and asserts the counts it printed, the
code it exited with, and the sentence it closed on — including that a run without the
flag does not blame it. Both come from one rule, gateFailed(), so the sentence cannot
describe a different rule from the one that produced the code:
$ php artisan db:doctor
Connection inspected: pgsql, database.default is pgsql
PASS provider swap WeightedDatabaseServiceProvider is in charge, weighted manager bound
PASS weighted factory db.factory is Uak35\WeightedDbManager\Database\Weighted\WeightedConnectionFactory
PASS published config config/db-manager.php is published and differs from the sample
FAIL pgcat gate swrr.pgcat.enabled is true but pgcat will never act: connection "sqlite"
uses driver "sqlite", and pgcat only fronts PostgreSQL. ... — recorded
unresolved since 2026-09-20T08:15:00+00:00
suggestion swrr.pgcat.enabled = false
PASS pgcat files flipper is not armed — no pgcat file is read or written
PASS pgcat supervisor flipper is not armed — no supervisor command is ever run
WARN replica metadata no read replicas on [pgsql] — every read uses the writer
PASS switch values swrr.pgcat.enabled on, swrr.pgcat.use_reload on,
swrr.allow_local_fallback on — each reads as on or off, so nothing
is refused
FAIL reader windows reader_windows[0] is "10:00-14:20" — each window must be an
array like ['start' => '10:00:00', 'end' => '14:20:00'] — a string
such as '10:00-14:20' is not a window. Refused: nothing is left
to apply, so reads use the replica pool at every hour
suggestion swrr.reader_windows = [['start' => '10:00:00', 'end' => '14:20:00']]
PASS store probe at most one PING every 60 seconds, throttled by the record at
/tmp/swrr-audit.json, which is writable
FAIL store reachability [redis(default)] could not serve a read: A connection attempt failed ...
[tcp://10.0.0.7:6379] — reads fall back to an in-process rotation that is
not shared between workers
11 checks: 7 passed, 1 warning, 3 failed
This installation will not behave as configured.
A FAIL on store reachability is the one worth reading twice: the application
still serves reads (they fall back), so nothing else looks wrong — but replica
choice stops being shared between workers, and the WARN form means this process
already fell back, which may predate the deploy. The row names which of the two
ways the store is unusable, because they are repaired differently: a connection
that refuses is Redis' problem, while the container has no "redis" binding yet
means the application never registered a Redis provider.
store probe exists because the reachability check can be silently unable to run,
and a skipped check looks exactly like a passing one. The probe is throttled by the
audit record, so an installation that cannot write that record — or that switched
probing off with store_probe_seconds = 0 — never issues a PING at all: no store
finding is produced, none is resolved, and the boot audit has nothing to say. The row
reports the interval, the record, and whether the record is writable, so an
installation that can never check its store is visible rather than merely quiet; the
reachability row's PASS in that state means "this boot did not check", which is why
this one sits directly above it. Disabling the probe is a choice, so that is a
warning; an unwritable record is not, because nothing then performs the check.
pgcat gate is the row that is not only this process's opinion. The boot audit
records a mismatch while it stands and clears it on the boot that finds the gate
open, so the row reports both halves: the verdict it computes now, and
recorded unresolved since … — the first sighting, preserved across boots — which
is what separates "this deploy is configured wrong" from "production has been in
this state since 2026-09-20". When the record names a connection this run does not
inspect, the row warns instead of passing silently, and --strict fails it.
replica metadata is the row whose question is easiest to get wrong. It is not "does
this installation look weighted" — a replica with 'weight' => 'heavy' looks weighted
and is not — but "is the pool the read list describes". The pool is what the resolver
kept, and the resolver reads a weight through max(0, ConfigValue::int(…)) and a core
count through max(1, …), so a value that is not a number becomes 0 — which is exactly
how a replica is switched off. An installation whose weight was a typo therefore lost
that replica from the pool without a word, and the row used to report the pool that was
left: a count short by one, with nothing to say which replica went or why. It now reads
the configured replicas and fails, naming each one and the value:
FAIL replica metadata replica metadata the resolver cannot read on [pgsql]:
[10.1.0.2:5432] weight is "heavy", which the resolver reads as 0
— a replica weighted 0 leaves the pool, so reads are routed over a
smaller pool than the read list describes and nothing else reports
that it happened
A cores or memory value is the same failure one key over, with the difference stated: the
replica stays in the pool, weighted as something other than the read list says,
because the resolver substitutes rather than drops it (cpu_cores becomes one core,
ram_gb none). The nearest case the rule has to get exactly right is weight: 0 itself,
which is the documented disable: that replica is meant to be absent, so the row passes —
and names it, because a pool reported as "1 replica, total weight 10" on an installation
with two is the same quiet shrink through the front door. The rule is the resolver's own
arithmetic rather than a restatement of it, which is what makes a negative weight, a
weight: 0.5 that truncates to 0, and a weight: null all reportable;
docs/replica-metadata-refusal.md records which values are refused, which are named, and
why weight: true is left alone.
reader windows answers a configuration typo rather than a configuration gap. An entry
that is not an array — the flat '10:00-14:20' — used to be dropped in silence, which
reads as no window at all: the resolver is then permissive, so every read uses the
replica pool, the opposite of the fallback the setting describes, and nothing said so.
The row fails on it, names the entry (reader_windows[0] is "10:00-14:20") and quotes
the shape a window has, then says what the resolver does next. Well-formed entries beside
a refused one still apply, so a typo costs the entry rather than the whole fallback. The
three remaining ways the setting can fail to act stay warnings, matching the boot audit,
because those are understood values that can do nothing — --strict fails them all the
same.
The same row answers the days setting, because it is the same mistake one key over:
swrr.reader_days = '1,2,3' is a list written as one string, and it used to be dropped
entry by entry until no day was left — which makes the resolver permissive, so a day list
meant to restrict reads runs all week as its opposite. The row fails on it and names the
value (swrr.reader_days is "1,2,3", not a list of days), quotes the shape the setting
reads, and says what is left: a list whose remaining entries are days keeps them, and the
row says how many.
One row can be about more than one of these at once, so it names every problem it finds
and not the first. The two settings are refused independently, and an operator who fixes the
windows and redeploys should not hear about the days on the next preflight — a second deploy
for a sentence that would have fitted on this row — while the boot audit had been reporting
both from the beginning. The verdict is the loudest problem in the list, so a refusal beside
a warning is still a FAIL, and each problem is dated from its own finding: the record
holds one entry per finding key, and quoting the first key on file would put a date on a
problem the row is not reporting. Two refused settings therefore print two suggestion
lines, one repair each.
Both failures end with a line that is not a check. A refused value that names its own
replacement is printed in the shape the setting reads, so the repair is a paste rather
than a translation — swrr.reader_windows = [['start' => '10:00:00', 'end' => '14:20:00']]
for the flat string, swrr.reader_days = [1, 2, 3] for '1,2,3' — and the row's FAIL,
the counts and the exit code are the same whether it is printed or not. A well-formed entry
beside the refused one is repeated exactly as it was written, extra keys and all, since that
is the value the setting already accepts. Where the refused value does not reduce to one
replacement there is no line at all: an overnight range ('22:00-06:00') has no equivalent
this resolver could enter, '9:00-14:20' has a bound the resolver silently reads as
midnight, '10:00 to 14:20' is not a range, and 'mon' is not a day. The package explains
the shape and stops, because a guess printed as a fix is worse than the sentence it
replaces. ReaderWindows::suggestion() and ReaderDays::suggestion() hold that rule;
docs/reader-windows-refusal.md lists which values get a line and why the rest do not.
The pgcat rows follow the same rule, and they are where it is easiest to see why it is a rule
rather than a nicety. pgcat gate prints swrr.pgcat.enabled = false when the switch is armed
where pgcat cannot act: the repair its own sentence already names, in the shape a report can be
acted on without being read as English — and, in --json, a string a gate can select on rather
than a clause buried in a detail. pgcat supervisor prints the setting line for the two faults
that reduce to a command, an unquoted program name and a command left empty. pgcat files prints
no line at all: every problem it reports is a path or a permission, the right value is whatever
this installation's pgcat and supervisor actually use, and a plausible-looking path pasted into
configuration replaces the operator's intent with this tool's. PgcatConfigFlipper::suggestionForGate()
and suggestionForSupervisor() hold that half of the rule, so the line comes from the class that
knows which setting the flip reads — and a FAIL stays a FAIL whether a line is printed under
it or not.
A FAIL on pgcat files is the one that would otherwise wait for the window to
open: the application boots, pgcat gate passes, and the flip throws at the moment
traffic was supposed to switch. Seeing it here means the flip could be repaired
before that moment rather than during it.
pgcat supervisor is the same failure one step later. pgcat files proves the file a
swap writes is reachable; the swap also runs a shell command, and everything about
that command is invisible to a file check — whether the executable resolves for the
user the flipper runs as, whether the program name reaches supervisorctl unchanged,
and whether supervisord knows that name. All four ways it can fail leave the file
already replaced:
FAIL pgcat supervisor a flip would run supervisorctl signal HUP pgcat:*, but pgcat:*
is unquoted: the shell may expand it before supervisorctl sees
it. Write supervisorctl signal HUP "pgcat:*" — the quoted name
is the one supervisorctl receives unchanged
suggestion swrr.pgcat.reload_command = 'supervisorctl signal HUP "pgcat:*"'
FAIL pgcat supervisor supervisor does not know "pgcat:*": pgcat:*: ERROR (no such
group). The name has to match what supervisord runs: a group
called pgcat is addressed as "pgcat:*", a single program by
its own name — a flip refuses before it swaps the file, so
nothing is replaced and no mode is recorded
The first of those two carries a suggestion and the second does not, which is the whole rule
in one screen: the unquoted name reduces to a command the package can write down exactly (and
names the key a flip reads — reload_command here, because use_reload is on), while "the
program supervisor does not know" is repaired in supervisord's own [program:] sections. The
row states the fault either way, and the file is not swapped either way.
The row asks supervisor rather than guessing, and asks it read-only: the command it runs
is the flip's own with status in the verb position, supervisorctl status "pgcat:*",
so nothing is restarted or signalled. The command itself comes from
PgcatConfigFlipper::supervisorCommand(), the method swap() uses, so the row cannot
describe a different command than the flip performs — including which of the two it
picks: the reload command when use_reload is true, the restart command otherwise.
A command that does not invoke supervisorctl (systemd, a wrapper script) is checked for
resolution only, and the row says so rather than implying it asked supervisor about
something. Like pgcat files, it is moot while the flipper is not armed.
The verdict behind all of this is one object — Pgcat\SupervisorStep::inspect() — which
the row, the flip's preflight and --dry-run all read, so a row that passes while a flip
refuses is impossible, and the sentence an operator reads is the same one in all three
places. The decision behind the read-only derivation, the quoting rule, the four ways the
step can fail, and what a flip now does about them is written up in
docs/pgcat-supervisor-preflight.md.
The preflight as data: db:doctor --json
A deploy gate should not parse the table either. --json writes one object and nothing
else — every row as checks, each with the verdict it reached and the repairs it carries,
the run's counts, the verdict for the run, and the exit_code the process exits with —
so the whole of stdout is the report. It is the same flag, envelope and rule as
db:pgcat-flip --json, so the package has one machine-readable contract rather than one per
command:
$ php artisan db:doctor --strict --json
{
"command": "db:doctor",
"connection": "pgsql",
"default_connection": "pgsql",
"strict": true,
"verdict": "FAIL",
"exit_code": 1,
"counts": {
"checks": 11,
"passed": 7,
"warnings": 1,
"failed": 3
},
"checks": [
{
"name": "provider swap",
"verdict": "PASS",
"detail": "WeightedDatabaseServiceProvider is in charge, weighted manager bound",
"suggestions": []
},
{
"name": "weighted factory",
"verdict": "PASS",
"detail": "db.factory is Uak35\\WeightedDbManager\\Database\\Weighted\\WeightedConnectionFactory",
"suggestions": []
},
{
"name": "published config",
"verdict": "PASS",
"detail": "config/db-manager.php is published and differs from the sample",
"suggestions": []
},
{
"name": "pgcat gate",
"verdict": "FAIL",
"detail": "swrr.pgcat.enabled is true but pgcat will never act: connection \"sqlite\" uses driver \"sqlite\", and pgcat only fronts PostgreSQL. ... — recorded unresolved since 2026-09-20T08:15:00+00:00",
"suggestions": ["swrr.pgcat.enabled = false"]
},
{
"name": "pgcat files",
"verdict": "PASS",
"detail": "swrr.pgcat.config_path is readable; swrr.pgcat.readers_path is readable; swrr.pgcat.no_readers_path is readable",
"suggestions": []
},
{
"name": "pgcat supervisor",
"verdict": "PASS",
"detail": "supervisorctl resolves to /usr/bin/supervisorctl and knows \"pgcat:*\"",
"suggestions": []
},
{
"name": "replica metadata",
"verdict": "WARN",
"detail": "no read replicas on [pgsql] — every read uses the writer",
"suggestions": []
},
{
"name": "switch values",
"verdict": "PASS",
"detail": "swrr.pgcat.enabled on, swrr.pgcat.use_reload on, swrr.allow_local_fallback on — each reads as on or off, so nothing is refused",
"suggestions": []
},
{
"name": "reader windows",
"verdict": "FAIL",
"detail": "reader_windows[0] is \"10:00-14:20\" — each window must be an array like ['start' => '10:00:00', 'end' => '14:20:00'] — a string such as '10:00-14:20' is not a window. Refused: nothing is left to apply, so reads use the replica pool at every hour",
"suggestions": [
"swrr.reader_windows = [['start' => '10:00:00', 'end' => '14:20:00']]"
]
},
{
"name": "store probe",
"verdict": "PASS",
"detail": "at most one PING every 60 seconds, throttled by the record at /tmp/swrr-audit.json, which is writable",
"suggestions": []
},
{
"name": "store reachability",
"verdict": "FAIL",
"detail": "[redis(default)] could not serve a read: Connection refused [tcp://10.0.0.7:6379] — reads fall back to an in-process rotation that is not shared between workers",
"suggestions": []
}
]
}
The report carries the rows in the order the table prints them, and detail is the row's own
detail — the same string, unwrapped rather than folded into the terminal's width. The pgcat gate
detail is abbreviated above with a trailing ...; a real one is several sentences long. The key set
is fixed and always present — an empty list where a run has nothing for a key — so a rule never has to
check whether a field exists:
php artisan db:doctor --strict --json > preflight.json jq -e '.exit_code == 0' preflight.json jq -e '.counts.failed == 0' preflight.json jq -e '.checks[] | select(.name == "store probe") | .verdict == "PASS"' preflight.json jq -e '.checks[] | select(.name == "reader windows") | .suggestions[]' preflight.json
verdict is one of three strings, at both scopes — the same three the table prints in its left
column, rather than a second vocabulary for the same fact:
| verdict | what it means |
|---|---|
PASS |
the check passed — and, for the run's own verdict, so did every row |
WARN |
the check found a setting that cannot act but is not refused — and, for the run, nothing failed and something warned |
FAIL |
the check refused a value — and, for the run, a row failed |
The run's verdict is the loudest row rather than a second opinion about the run, and it is
deliberately not read back out of exit_code: under --strict a warning exits 1 while every row is
PASS or WARN. So a gate that wants to know whether the installation is broken asserts on
verdict, one that wants the gate's decision asserts on exit_code, and the two fields that join
them are strict and counts:
jq -e '.verdict != "FAIL"' preflight.json # is anything actually broken? jq -e '.exit_code == 1 and .verdict == "WARN" and .strict' preflight.json # ...or is this just --strict?
Two things to know about the row list. It is not a fixed length: a db that does not resolve to the
weighted manager leaves the six configuration rows and no replica rows at all, so select a row by
name — the row's identity, and the same string the table prints — rather than by position. And
detail is prose written for a reader, naming paths and commands and supervisor's own output on the
platform that ran it; assert on name, verdict and exit_code, which are the contract, and treat
suggestions as the repairs a row offers rather than as a substitute for reading the row.
--json changes the report and nothing else: the same checks run, the code is the one gateFailed()
computed, and nothing is written either way. It composes with --strict, which is how a pipeline will
pass it, and there is no combination to refuse — a doctor is not a daemon, so a report cannot be
untrue of the run that produced it. The verdicts above are read back out of this table by
DbDoctorTest::test_every_json_verdict_is_documented, the same guard the flip's kind table has, so a
verdict cannot be introduced without documenting it or documented without something producing it. The
decision — including the three candidates it shares with the flip's record, and the two that are the
doctor's own, a rows-free envelope and a per-row exit code — is in
docs/db-doctor-json.md.
Flipping: db:pgcat-flip
The flip's exit code is the same kind of contract as the sweep's and the preflight's — it is
what a scheduler, or a --dry-run gate on a deploy, branches on — so it is written the same
way:
| what the run found | exit |
|---|---|
| a flip applied, nothing to do, or skipped by another instance | 0 |
| a flip that was forced | 0 |
| a rehearsal, whether it would flip, would not, or was forced | 0 |
| the flipper is not armed for this connection | 0 |
--status: a state report was asked for |
0 |
| a step a flip needs did not work, or a rehearsal of one | 1 |
| a flag combination that is refused | 1 |
| the container has no pgcat flipper | 1 |
The rule behind the table is one sentence: non-zero exactly when the run could not do what it
was asked. Applying a flip, having nothing to apply, and being skipped by another instance are
all the run doing its job, so a scheduler never has to read the output to tell them apart —
whereas a step that did not work, or a request that was refused before it started, is a run that
did not happen. The two refusals are --force-mode=sideways, which is not a mode, and
--dry-run --watch, which would be a daemon whose every pass is a rehearsal.
Every row is pinned by
DbFlipPgcatCommandTest::test_the_exit_code_is_a_function_of_what_the_flip_would_do, which
builds each state as a real installation in a temp directory and asserts the code, the line
beside it, the file it did or did not touch, and whether the run recorded a mode. The table is
then read back out of the README by
DbFlipPgcatCommandTest::test_the_matrix_agrees_with_the_readme_exit_table, the same guard the
sweep and the preflight have — so a case cannot be documented without a cell behind it, and a
cell cannot be added without documenting it.
One state the table names is not a row of that matrix and has a test of its own: --status
wins when a rehearsal is asked for in the same run, because its only evidence is the line the
other branch did not print. What each flag does, and which half of a flip a rehearsal can
prove, is under Pgcat.
The JSON report: one object for a pipeline
A pipeline should not parse a rendered report. --json writes one object and nothing else — the
verdict, the code the process exits with, and the evidence of whichever route produced it — so
the whole of stdout is the report:
$ php artisan db:pgcat-flip --dry-run --json
{
"command": "db:pgcat-flip",
"kind": "would_flip",
"exit_code": 0,
"mode": "readers",
"previous_mode": null,
"reason": "the mode has never been applied, so a first flip would run",
"error": null,
"steps": [
{"step": "lock", "outcome": "done", "detail": "taken and released: /…/pgcat-flip.lock"},
{"step": "read source", "outcome": "done", "detail": "/…/pgcat-readers.toml (35 bytes)"},
{"step": "supervisor check", "outcome": "done", "detail": "supervisorctl resolves to /…/supervisorctl and knows \"pgcat:*\": pgcat:pgcat_00 RUNNING pid 4242, uptime 0:12:34"},
{"step": "write temp", "outcome": "done", "detail": "/…/pgcat.toml.tmp.47388 (35 bytes)"},
{"step": "remove temp", "outcome": "done", "detail": "the rehearsal leaves nothing behind"},
{"step": "rename", "outcome": "would", "detail": "/…/pgcat.toml.tmp.47388 → /…/pgcat.toml (the content would change)"},
{"step": "supervisor", "outcome": "would", "detail": "/…/supervisorctl restart \"pgcat:*\""},
{"step": "state file", "outcome": "would", "detail": "/…/pgcat-flip-state.json (writable, left as it is)"}
],
"status": null
}
The report pretty-prints each step over several lines and names real paths; they are shown one
per line here. The key set is fixed and always present — null, or an empty list, where a route
has nothing for a key — so a rule never has to check whether a field exists:
php artisan db:pgcat-flip --dry-run --json > rehearsal.json jq -e '.exit_code == 0 and .kind == "would_flip"' rehearsal.json jq -e '[.steps[].outcome] | all(. == "done" or . == "would")' rehearsal.json
kind is the verdict, from one vocabulary:
kind |
what it means | exit |
|---|---|---|
flipped |
a flip applied | 0 |
no_change |
the mode on record already matched, so nothing was done | 0 |
skipped |
another flipper instance held the lock | 0 |
disabled |
flipping is off, or pgcat cannot front this driver | 0 |
status |
a state report was asked for (--status) |
0 |
would_flip |
a rehearsal proved a flip would apply | 0 |
would_not_flip |
a rehearsal found nothing to do | 0 |
failed |
a step a flip needs did not work — a flip, or a rehearsal of one | 1 |
refused |
a flag combination was refused before anything was read | 1 |
unbound |
the container has no pgcat flipper | 1 |
mode and previous_mode are the modes the run is about. A status report carries its two
inside status, which is the flipper's own status() array — with null for an unset path
rather than the (not set) the table substitutes, because a job has to be able to test a path
for being unset. steps is a rehearsal's pipeline and empty for every other route; reason and
error are the sentence a refusal or a failure would otherwise have printed down a human
channel.
Two things --json does not do. It does not move a single exit code —
DbFlipPgcatCommandTest::test_the_json_report_is_the_same_run_as_one_object runs this whole exit
matrix twice, once for the rendered report and once for the object, and asserts the code, the
file and the record are the same either way. And it cannot be combined with --watch: a report a
machine reads is one object, and a daemon would print one every interval without ever reporting a
verdict for the run. The kinds above are read back out of this table by
DbFlipPgcatCommandTest::test_every_json_kind_is_documented_with_its_exit_code — the same guard
the exit tables have, so a kind cannot be introduced without documenting it or documented without
something producing it.
Reporting failures from your own code
try { DB::select('SELECT …'); } catch (\Illuminate\Database\QueryException $e) { app('db')->markReplicaFailed('10.0.1.11', 5432, $e->getMessage()); }
The replica is excluded from routing for an exponentially growing cool-down (30 s, 60 s, 120 s … capped at 30 min) and retried afterwards.
Layout
src/Database/Weighted/ SWRR algorithm, weight resolver, health monitor, state stores,
connection factory, window resolver, manager
src/Pgcat/ pgcat.toml flipper, the supervisor step it runs afterwards,
and the two result value objects (a flip, and a rehearsal)
src/Providers/ WeightedDatabaseServiceProvider (replaces the framework's)
src/Console/Commands/ db:doctor, db:replica-status, db:probe-replicas, db:pgcat-flip
src/Http/Controllers/ DatabaseHealthController
src/Support/ ConfigValue (typed config reads), ReaderWindows/ReaderDays (what the
reader fallback may contain), BootAudit — findings and their record
config/db-manager.php the sample config — `vendor:publish` copies it into the app,
which then owns and edits its own copy
config/app.php Laravel 12 skeleton + the provider-swap block above
tests/ unit tests plus a Testbench test that boots the provider
tests/Support/ fixtures the command tests build their rows from — a pgcat
installation in a temp directory, a stand-in supervisor, the
real manager with a probe's reporting calls recorded, and the
README read as data, so a documented exit code is asserted
rather than merely written down
docs/ design records for decisions that are not obvious from the
code — the boot audit's store-probe timing, where its findings
are published, and what it does when two of them claim one key,
what the package does with a reader window that is not a
window, how the flip's supervisor step is checked without
taking it, which half of a flip can be rehearsed and how a
pipeline reads its verdict, what a deploy gate reads out of the
preflight, what the preflight does with replica metadata it
cannot read, and how the documented exit codes are kept in step
with the code
bin/checks.php every local check, summed up in one summary
bin/release.php tag-driven releases (see RELEASING.md)
bin/surface.php the symbol reader and differ both release commands share
bin/inventory.php writes files.tsv and methods.tsv from the working tree
bin/publish-config.php publishes the sample config into an application
files.tsv, methods.tsv what the last release shipped — its files and their public
methods, stamped with the tag they describe, refreshed by
bin/release.php and read back by the next release
Testing
composer install composer test # the three gates CI runs: pint --test, phpstan, phpunit composer checks # everything that can be checked, in one summary vendor/bin/phpunit --testdox
composer checks (bin/checks.php) runs every applicable check and prints one
summary with one exit code: php -l over every file (including bin/), an
independent AST parse of the same files, composer validate --strict,
check-platform-reqs, the workflow YAML, PHPStan, Pint and PHPUnit. It prints only failures by default (--verbose streams everything), takes
--only=syntax,tests to narrow, reports
a missing tool as a skip rather than crashing on it (--require-all turns that
into a failure), and can add composer audit with --audit.
phpstan.neon.dist runs at level max over src with no ignore entries
and no baseline file — every finding is fixed, not silenced.
STATIC-ANALYSIS.md records the staged burn-down from
level 6, the per-level numbers and the behaviour changes that came with it.
Releasing
Releases are cut from git tags — there is no version field in composer.json
and no version file, so a tag is the only version that can be out of date:
composer release -- --weigh --dry-run # print the plan, change nothing composer release -- --weigh # promote the CHANGELOG, commit, tag composer release -- --weigh --push # ...and push branch + tag (or push yourself)
bin/release.php works the next version out from the latest tag (the base is
0.0.0 when there are none) and the bump out of the changes themselves:
--weigh reads four signals — the Unreleased notes, the commits since the last
tag, the public surface of src/ and config/, and the inventory (files.tsv,
methods.tsv) the previous release wrote — and applies the policy in
RELEASING.md — a fix is a patch, a new command or config key is a
minor, a removed or renamed public symbol is a minor while 0.x and a major after
1.0.0. Declaring a smaller bump instead (--minor, --major,
--version=X.Y.Z) stops the release with exit code 1 and names the change that
forbids it, so a breaking change cannot be shipped as a patch. It then promotes
## Unreleased to the released heading with a compare link, refreshes both
inventory files in the same commit, keeps both dev-lane branch aliases
(dev-main, dev-dev) on the line being developed, and refuses to continue on a
dirty tree, the wrong branch, an existing tag or an empty release section. A
version may also be a prerelease — --version=0.0.1-alpha1, cut from dev — which
promotes the notes and stamps the inventory exactly as a release does; a consumer
opts in with minimum-stability: alpha or a constraint like "^0.0.1@alpha". The
suffixes it writes are the ones Composer's own parser reads, and the spellings it
cannot read (0.0.1-dev.1, 0.0.1-alpha.1) are refused rather than written.
RELEASING.md has the policy, the safety rails and the first-release
walkthrough.
Changelog
The five defects this package was ported with — unrunnable routing, an unreadable
formula key, disagreeing factor defaults, a broken db:probe-replicas, a
hardcoded local-store prefix and silent degradation — are documented with their
fixes in CHANGELOG.md.
License
MIT — see LICENSE.