Picture the back rack of a fastener godown. Steel shelving, six feet of
bagged bolts on either side, a handset in one hand and a barcode gun in
the other. The signal indicator shows one bar, then none, then one again.
Thirty bags have just come off a lorry and need to be booked in before
anyone can sell them.
Most stock software was not built for this moment. It was built in an
office, tested on office Wi-Fi, and it treats the network as a given and
its absence as an error. Expecting a steady connection at the back of a
godown is quixotic. The software has to expect the opposite.
What follows are the rules we arrived at building the
Lucky Traders Stock app, a Windows and
Android application that runs a fastener godown. Some of them we knew on
day one. Some we learned in the first weeks of live use.
1. Queue every write, always
The obvious design is to send a change straight to the server when there
is a connection and queue it only when there is not. The better design is
to queue every write, every time, and let a background flush deliver it.
Two reasons. A picker should not wait on a round trip per carton, even on
a good connection. And "there is a connection" is a guess that is wrong
at the worst moment: the request leaves, the signal drops, and the app
has to decide what happened.
The lesson here came the hard way. The first versions queued only the two
writes everyone thinks of, scans and deliveries. The other edits (a
renamed item, a reorder level, a vendor's phone number, a rack count) went
straight to the server and failed silently without a signal. The app
opened offline and then appeared not to save anything. Every write path
now goes through a queue. Not two of them. All of them.
2. Make every write safe to send twice
A request can commit on the server and then time out on the way back. The
device sees a failure and retries. Without protection, the delivery is
booked in twice.
The fix is an identifier minted on the device for every queued write. The
server records it. A retry carrying the same identifier is recognised and
answered, not applied again. For a delivery of thirty bags, each bag
derives its own key from the delivery's key, so a replayed lorry cannot
double the stock.
3. Mint identifiers on the device
A sticker has to come off the printer while the bag is still on the floor.
If the sticker's code is issued by the server, it cannot exist until the
server has heard about the delivery, and with no signal, the bags wait.
So the device generates the codes itself, in a format long enough that a
collision is vanishingly unlikely, and the database enforces uniqueness
in case one ever happens. The same applies to anything created offline: a
new rack gets a temporary local name, and when the server answers with
the real identifier, the app swaps it through every queued entry that
refers to it.
4. A refusal settles the entry, it does not block the queue
Sometimes the server says no, for a reason that will never change. Somebody
else deleted that rack. That item name now collides. If the refused entry
stays at the front of the queue, every later change on that device is
stuck behind it forever.
So a refusal is recorded, shown to the person in plain words, and taken
out of the way. Losing one edit is bad. Losing every edit behind it is far
worse.
5. Cache who is signed in
This one cost a whole feature for a while. The app cached the item list,
the racks and the warehouses, but not the session. Started with no signal,
it asked the server who was signed in, got no answer, and dropped to the
sign-in screen, where it could not reach a single figure it was already
holding. It looked like a login problem. It was a caching problem.
The cached session grants nothing on its own. The device token is what
authorises, it lives in the operating system's credential store, and the
server checks permission again on every request. The cache only decides
which buttons to draw.
6. Date the movement, not the upload
A queue that sits overnight and flushes in the morning must not book
yesterday's scans as today's. The date a bag physically moved travels with
each queued entry, and a flush holding two days of scans goes up as two
requests, one per day.
The default date is set by the server, in Indian Standard Time, not by the
handset. A phone with a wrong clock would otherwise date a whole delivery
wrongly, and nothing would look amiss until somebody read a report.
7. Say plainly what does not work offline
Some things should not work without a connection. Reorder suggestions,
purchase orders and reports are computed from figures that move every
time anybody scans anything. A cached answer would be confidently wrong,
and confident wrongness is the most expensive kind.
So those screens say, in words, that they need a connection. Everything
else keeps working. A status line under the header is always visible:
online or offline, when the device last synced, and how many entries are
waiting to send. People trust a system that tells them what it does not
know.
The test that matters
Offline-first sounds like an engineering preference. It is an operational
one. The test is simple: walk to the furthest rack in your godown, switch
the phone to airplane mode, and try to receive a delivery, print its
stickers, dispatch a bag and correct a mistake. Then switch it back and
check that every one of those landed, once, on the right day.
If your current stock software fails that walk, the problem is not the
signal. The signal was never going to improve. Start a conversation about
software that assumes it won't.