Home / Blog / Developers & integrators
Developers & integrators

Aqara Push Subscription and Webhook Design for Production

Server room corridor with event streams visualised as lines of data flowing between connected smart home devices

Answer up front — Aqara's message push service has two retrieval routes: HTTP push to an address you configure, and MQ, described in the docs as built on open-source RocketMQ. The HTTP route is verified periodically, failure statistics are kept, and the enforcement rule is real: if the push failure rate exceeds 5% within 5 minutes you are notified by SMS or email, and half an hour after that notification the push is suspended until you re-enable it in the console. Both routes retain messages for the last 12 hours. Every message carries a msgId and millisecond timestamps — build idempotency and reconciliation on those, and treat ordering as something you handle rather than assume.

This is the post that stops your integration from quietly failing. A push integration that is ninety-nine percent reliable looks perfect in a demo and corrupts a customer's data in month two.

The two channels, and when each earns its place

The documentation is direct about the two methods and the reason for both: meeting requirements for message real-time and message persistence.

HTTP pushMQ
ShapePlatform POSTs JSON to an address you registerPlatform publishes to a queue; you subscribe and consume
Retained windowLast 12 hours of failed pushes, queryable by APILast 12 hours of messages kept in MQ for consumption
Delivery couplingYour endpoint must be up and fastYour consumer acknowledges
Failure visibilityThe platform polls your address periodically and tracks failure rateNot described in the English docs
Best fitA normal web service that can be scaled horizontallyHigh-volume fleets where you want the buffer

HTTP is the default choice for most integrators. Reach for MQ when your per-device event volume makes the 12-hour HTTP failure window too thin a safety net, or when your consumer is already a queue consumer. Note the asymmetry: for HTTP the docs describe an explicit verification and failure-tracking mechanism; for MQ they describe retention and subscription, but no equivalent enforcement rule. If you choose MQ for reliability, do not assume it self-heals the same way.

What the platform documents — and what it does not

Stated facts, from the message push pages:

  • Delivery is HTTP POST, application/json, to the address configured under Console → Project Management → Message push settings.
  • Required headers: token (a valid access token for the authorised account), time, and nonce (random, for uniqueness). Optional: appkey and sign, returned only if you enable signature verification with an appKey and appSecret on that page.
  • Push signature, when enabled: sort appkey, nonce, token, time by ASCII → splice as appkey=xxx&nonce=xxx&time=xxx&token=xxx → append the appSecret → lowercase → MD5 32-bit.
  • The address is verified from time to time to ensure the reliability of the service address and the message-receiving response mechanism.
  • Failure threshold: above 5% failure rate within 5 minutes, the third party is notified by SMS or email. If not resolved, the push is suspended half an hour after the notification. Re-enable it in the console.
  • Failed push messages are kept for the last 12 hours, with a query API provided for them.

Not stated in the English docs, and worth verifying with Aqara before relying on them: at-least-once versus at-most-once, ordering guarantees between devices, retry policy and retry count for a failed POST, and whether a non-2xx counts toward the 5% statistic. Design as though every message can be duplicated, delayed, reordered and lost.

Idempotency: assume the repeat

Device state is level-triggered. A door sensor reports open; it does not report an edge. A resend of the same state is not an error, and a consumer that treats every message as new will double-count everything.

The platform gives you what you need. Every message carries msgId, a message unique identification id; time, the timestamp when the message was generated, in milliseconds; openId, the authorised user identifier; eventType; and data, with its own data.time in milliseconds. Store msgId with a unique constraint and drop duplicates on arrival, at the edge of your consumer — and put the insert and the state write in the same transaction, because a dedupe table written outside it will lie to you.

Three message categories arrive, and they need different handling:

  1. Event notification messages — device lifecycle facts: bind and unbind (gateway_bind, subdevice_bind, gateway_unbind, unbind_sub_gw), online and offline (gateway_online, gateway_offline, subdevice_online, subdevice_offline), dev_name_change, dev_position_assign, and the rule events linkage_created, scene_created, event_created and their _deleted counterparts. The docs state these are all pushed to third-party servers — you cannot filter them.
  2. Device attribute messages — state changes and operation triggers such as switch state, load power and power consumption, pushed according to the subscription mode the user selected.
  3. Device control failure messages — returned when control fails, with the trigger source, trigger time and an error code. Your only reliable signal that a control you sent did not take effect.

Subscribing: narrow it, or drown in it

The docs are explicit that device data volume is huge, and that you can narrow what you receive. Two subscription interfaces, both requiring message push to be configured first:

  • config.resource.subscribe — a list of resources, each with a subjectId, an array of resourceIds and an optional attach string.
  • spec.config.trait.subscribe — an array of traits, each with a deviceId, an array of codePaths in the documented format endpointId.functionCode.traitCode, and an optional attach.

The attach field is worth using even if you do not need it. It passes through transparently into the notification body, which makes it a natural place for your own correlation data — which subscription, which tenant, which logical device. Subscribe per tenant, not globally: a 2,000-unit estate where every consumer receives every attribute message is a cost and latency problem you create for yourself.

Ordering: two clocks, neither trustworthy

Compare time against data.time. Both are millisecond timestamps and they are not the same thing — the outer one is when the message was generated, the inner one is the specific event's timestamp. That gap is where out-of-order delivery shows up: a motion sensor triggers, a second trigger arrives while the first is in flight, and a clear lands before an occupancy set. A consumer applying them in arrival order ends up with a room that believes it is occupied when it is empty. Without documented ordering guarantees:

  • Use the inner data.time for last-writer-wins, not arrival order. Keep current state alongside its timestamp and reject an update older than what you hold.
  • Tolerate reordering on transient attributes — presence, motion, occupancy. Set them as a boolean from the latest timestamp, never a counter and never a toggle.
  • Do not derive business state from the absence of a message. Nothing in the docs promises a timeout message. Absence of data is not data.

Reconciliation is the real answer

If you cannot get a guarantee, you do not need one — you need a repair loop. Schedule a periodic full-state read of the devices you care about and overwrite your local view. Everything the push stream got wrong while you were down, out of order or deduplicated incorrectly, a full read fixes.

Size it to bound the window of visible wrongness while staying affordable: a full sweep per space every few minutes is defensible for a building-management overlay, hourly is plenty for a consumer app. The 12-hour retained window for failed pushes is an asset here — have reconciliation consult that failure query first, since replaying what was dropped is cheaper than re-reading everything.

When your endpoint is down

The documented failure path is unambiguous and it has teeth: above a 5% failure rate within 5 minutes you are notified by SMS or email; half an hour after that notification, push is suspended until you re-enable it in the console. A twenty-minute deploy can leave your integration silently unsubscribed, with the only alert on a phone nobody is watching.

Defences, in order of value:

  • Return 2xx fast. Acknowledge, then process, on a queue behind the handler.
  • Alert on the platform's channel, not just your own. The SMS/email arrives before your suspension. Route it to whoever can act in minutes.
  • Monitor your own drop rate. A 5-minute error rate near 5% puts you in the enforcement window whether or not anyone noticed.
  • Automate the re-enable path or write the runbook before you need it. Suspended push is a console toggle, and at 2am that toggle is the whole incident.
  • Fail loudly in the product. A UI showing "last updated 4 hours ago" beats an alert nobody reads.

Pre-production checklist

  • [ ] Signature verification enabled and tested, or consciously left off
  • [ ] msgId unique constraint inside the same transaction as the state change
  • [ ] Subscriptions scoped per tenant, not global; attach used for correlation
  • [ ] Last-writer-wins keyed on data.time, not arrival order
  • [ ] Control failure messages routed where a human will see them
  • [ ] Periodic full-state reconciliation scheduled and costed
  • [ ] Failure-rate alert at a safe margin below 5% in 5 minutes
  • [ ] Runbook for re-enabling suspended push
  • [ ] Delivery guarantee, ordering, retry policy and rate limits confirmed with Aqara in writing

Planning a project?

Tell us about your space. Our B2B team will reply within one business day with a recommended setup and quotation.

WhatsApp us →
[email protected]
+603-5880 5486