Cover illustration for “Sending Slack Messages From an Agent Without Flooding a Channel”
Tool & App Connections

Sending Slack Messages From an Agent Without Flooding a Channel

September 30, 202611 min read2,373 words

Share

Text: Freya Whitfield

Event subscriptions replace polling to let agents read Slack without burning through rate limits.

Wire an agent up to Slack, and the first instinct is almost always the same: call chat.postMessage and let it talk. Every finding, every intermediate step, every small update gets posted as it happens. That instinct produces two kinds of failure, and only one of them is obvious.

The obvious one is social. Humans get tired of the noise and mute the channel. Fair enough, that's a UX problem, and most teams have run into some version of it. The second failure is operational, and it's the one nobody warns you about: the agent itself starts losing access to its own context, quietly, with no error message telling you why. The same rate-limit tightening that governs how much Slack history an agent can pull also determines whether that agent can stay coherent across a conversation at all.

Everything downstream in this piece, the polling failures, the event-based fix, the posting discipline, exists as a response to that constraint. None of it is optional styling anymore.

Slack's 2025–2026 rate limit changes for non-Marketplace agents

Slack didn't do this quietly. Effective May 29, 2025, Slack changed rate limits for conversations.history and conversations.replies for newly created non-Marketplace apps. Existing installations got a grace period. That period ran out on March 3, 2026, and from that date forward, every non-Marketplace app is subject to the new limits.

The actual numbers tell the story better than any description of them. Before the change, an app could call conversations.history or conversations.replies roughly 50 times per minute, and each call could return over 100 messages. After the change, once fully enforced, that drops to 1 request per minute, with a hard cap of 15 messages per response, a narrow throttle. That's a narrow throttle. That's a different API.

Who gets hit depends on distribution status. Apps already approved for the Slack Marketplace are exempt, and so are internal apps built by a company for its own workspace. The limit is aimed squarely at apps distributed outside the Marketplace that haven't gone through Marketplace approval. Except that API comes with the same restriction baked in: it's available only to directory-published apps and internal apps. An unlisted third-party app gets neither the old fast history access nor the new replacement. It's boxed out of both doors.

The rollout, for anyone trying to figure out why their app broke on a specific date, went in two waves. New apps and new installations felt it starting May 29, 2025, and net-new installations of already-existing apps were caught in that same wave. Everyone else, the apps that had been running fine for months or years, lost their exemption on March 3, 2026. Slack's stated reason is preventing data exfiltration by apps that haven't been vetted. Outside observers have pointed out the practical effect looks a lot like steering teams away from third-party LLM access to Slack data and toward Slack's own AI products. It's a small detail, not a central issue.

One more constraint belongs in this section, because it governs the other half of the problem: posting. chat.postMessage allows 1 message per second per channel, with short bursts sometimes tolerated but never guaranteed. Go over that and messages get silently dropped, or the user sees an error. There's a separate ceiling for events: 30,000 event deliveries per workspace per app per 60 minutes. That number becomes important once the conversation shifts from reading to architecture.

The standard agent polling pattern's breakdown under these constraints

Most agents built for Slack follow a pattern that made total sense under the old limits and falls apart under the new ones. On every incoming message, the agent calls conversations.history to grab recent channel context, then calls conversations.replies to read the thread. Two calls, every single interaction.

Under the old ceiling of roughly 50 requests per minute, returning 100+ messages each, that pattern was invisible. Fast, cheap, no one noticed it happening. Most agent frameworks aren't built to sit around for four minutes. They time out, or they just proceed with whatever partial context they've already got.

What that looks like in practice is unglamorous but real. The agent forgets anything that happened earlier in the day, because it can't see past that 15-message window. And when the framework does try to be clever and retries on an HTTP 429, response times start spiking at random, with no obvious pattern the user can point to.

There's a well-documented case from a Slack developer's own codebase that illustrates exactly this failure mode, even though it involved a different endpoint. A cleanup job was written to close old group DM conversations using conversations.close, and the AI-generated code looked completely reasonable: fetch the old conversations, loop through them, close each one. It ran fine in testing. Then it hit a workspace with hundreds of stale conversations, immediately slammed into Slack's 1-request-per-second limit on that endpoint, and because Slack's rate limits are global across all methods for a workspace, it took the entire application down. Every other call, posting, fetching user info, all of it, started failing at once. The fix that got proposed first was worse: catch the rate limit error, sleep for the retry-after duration, then retry, all inside the same synchronous loop. That blocks the entire process. In a system meant to be running concurrent operations, a blind sleep in one code path stalls everything else waiting behind it. The bug wasn't visible by reading the method in isolation. It only appeared once the system was running against real data at scale, which is exactly the trap with rate limits: they're a property of the whole system, not of any one function.

The second common mistake compounds the first. Some agents try to stay "helpful" by posting every intermediate output as soon as it's generated. That approaches the 1-per-second posting limit and floods the channel, training humans to mute the bot. Two different failure modes, same root cause: treating Slack's API like it has no limits worth designing around.

Switching from polling to event subscriptions as the foundational fix

The fix isn't a smarter retry loop. It's a different architecture entirely, and it starts with understanding an asymmetry that most agent builders miss on their first pass. The Web API's rate limits, the ones throttling conversations.history down to 1 call a minute, apply to pulling datac20.

Subscribe to message events, and Slack sends every message to a webhook the moment it's posted, in real time, with no polling involved and no rate limit standing between the agent and incoming messages, short of that 30,000-per-hour ceiling. That single shift changes the whole shape of the problem.

Instead of calling conversations.history after the fact to reconstruct what happened, the agent keeps its own running store of the conversation, a database, even something as simple as SQLite, that captures the message stream from the moment the subscription starts. Once that's in place, the agent needs zero Web API calls for context under normal operation, and it responds faster, because there's no round trip to Slack for history before it can answer.

It's not free of trade-offs. There's no fast way to backfill history from before the subscription began, and the whole approach depends on someone actually maintaining that local store with discipline. Fifteen messages of Web API history alone isn't enough to keep an agent coherent, but 15 messages layered on top of a running stack of conversation summaries is workable. That makes summarization a core piece of the architecture, not a feature to bolt on later.

Hermes connects over Socket Mode using the modern Bolt SDK, which uses WebSockets rather than a public HTTP endpoint. That means no public server requirement, and the agent can sit behind a firewall or on private infrastructure without issue. Slack deprecated the older RTM API for legacy custom bots back in March 2025, and full classic app deprecation is scheduled for November 2026. Socket Mode isn't just a nice option anymore; it's the direction things are heading.

Controlling what the agent posts

Fixing the read side solves half the problem. The other half is what the agent chooses to post, and this is where most of the actionable, low-effort fixes live.

Start with the simplest one: reply in thread instead of posting to the top of the channel. Thread-first requires no new infrastructure, no new code path, and it belongs in every agent that posts into a shared channel. It's the single cheapest intervention available.

From there, batch. During an alert storm, a single message reading "12 alerts firing across checkout" serves the rate limiter and the human reading it far better than 12 separate pings. Deduplicate next: if the same alert or output keeps firing, post it once and update a count rather than repeating the same message over and over.

Where possible, update in place with chat.update instead of posting fresh.

Then there's a judgment call that sits above all of these: significance gating. Before the agent calls chat.postMessage at all, it should ask whether this output actually clears a bar worth a human's attention. Intermediate reasoning, low-confidence guesses, routine status pings, none of that belongs in a shared channel. Put it in a thread, a DM, or just suppress it.

Last, treat the 1-per-minute read budget as something to spend deliberately. Keep a sliding-window counter of conversations.history calls per workspace, and only spend a call when the cache is genuinely cold, say, on a daemon restart or the first interaction in a channel. Save that budget for the moments where it actually matters, because there won't be many chances to use it otherwise.

Handling HTTP 429 responses correctly when the limit is hit anyway

Even a well-designed agent hits a 429 sometimes. What matters is what happens next, and Slack actually tells you exactly what to do: the 429 response includes a Retry-After header specifying the exact number of seconds to wait before trying again. That header is the API handing you the answer directly. Parse it, wait that long, then retry, no fixed sleep timers and no immediate retries.

Jitter matters more than it sounds like it should. If multiple agent instances all fail at the same instant and retry at the same interval, the result is a synchronized burst against the API, effectively a self-inflicted denial-of-service against Slack's rate limiter. Exponential backoff with randomized jitter spreads those retries out across a window instead of stacking them at one instant, which is the actual fix here, not a footnote to it.

Ignoring the Retry-After header, and either failing silently or retrying right away, is the common defect. Ignoring the Retry-After header entirely and either failing silently or retrying immediately are both worse than waiting.

One scoping detail that changes how urgent a 429 actually is: it only throttles that specific method for that specific workspace. A 429 is narrow, as long as the retry logic doesn't turn it into something bigger.

Connecting OpenClaw, Hermes, Claude Code, and Codex to Slack

The patterns above aren't theoretical. They map onto the runtimes people are actually deploying right now.

OpenClaw runs as a single long-running Node.js Gateway process that bridges several messaging channels, Slack, Discord, Telegram, WhatsApp, iMessage, Signal, Teams, and Matrix among them, to LLM backends. It registers as an external OAuth app, which puts it squarely in the non-Marketplace bucket and directly under the March 2026 enforcement. Its Agent Client Protocol can hand sub-tasks off to tools like Claude Code and Cursor, and every sub-task completion is a potential Slack post. Without batching those completions before they hit the channel, OpenClaw setups flood a channel fast.

Hermes connects over Socket Mode using the modern Bolt SDK, WebSocket-based, no public endpoint needed. Its default behavior already builds in the thread-first pattern: it replies in thread when @mentioned in a channel, with subsequent replies in that active thread continuing without a mention, and it responds to every DM without a mention. On the security side, tokens sit in ~/.hermes/.env with 600 permissions, and SLACK_ALLOWED_USERS has to be populated before the agent goes anywhere near a real channel, otherwise it denies every incoming message by default.

Claude Code is built for software development work, not general-purpose automation, and its Slack integration shows up as part of scheduled routines and background agent work in Anthropic's documentation. It runs on Claude models by default, though gateway proxies allow other models in. Its natural output pattern is reporting after a task finishes, which lines up directly with the significance-gating idea: post the result, skip the intermediate steps.

Codex, OpenAI's cloud-based coding agent, lists Slack as an integration alongside Linear, and shows up across the app, an IDE extension, the CLI, and the web. Its non-interactive mode and the Codex SDK enable automated workflows where Slack notification is a terminal action, and update-in-place and batch patterns apply when multiple workflow runs complete in close succession.

Codex uses AGENTS.md, Claude Code uses CLAUDE.md, and OpenClaw and Hermes support custom skill definitions (all accept plain-text behavioral instructions). Posting behavior, when to thread, when to batch, what never gets posted at all, can be written once into those files and carried across runtimes with barely any rewriting. That's a small thing to set up and it saves relitigating the same posting rules every time a new agent gets stood up.

Using a third-party tool-integration platform to manage Slack tool calls across runtimes without duplicating integration logic

Every pattern above, threading, batching, deduplication, backoff with jitter, significance gating, has to actually get implemented somewhere, and doing it separately for every runtime an agent might run on is its own kind of tax.

That matters more once an organization is running more than one of these agents at once, which is increasingly the normal case rather than the exception. A batching rule written once, a Retry-After handler written once, a significance threshold defined once, all of it stays consistent instead of drifting into slightly different behavior on each runtime. Given how easy it is for one team's OpenClaw instance and another team's Claude Code setup to end up with quietly different Slack behavior, having the integration logic centralized rather than copy-pasted across codebases is less a convenience and more a way of making sure the hard-won patterns from earlier in this piece actually stick.

Sources

  1. Rate limits - Slack Developer Docs
  2. AI Slop: A Slack API Rate Limiting Disaster – code.dblock.org | tech blog
  3. Rate limit changes for non-Marketplace apps | Slack Developer Docs
  4. Rate limit changes for non-Marketplace apps | Slack
  5. The Events API | Slack Developer Docs
  6. Handling Rate Limits with Slack APIs | by Shay DeWael | Slack Platform Blog | Medium

More in Tool & App Connections