Let's Automate My Social Life
Replying on KakaoTalk has always been hard. I wondered whether developing social skills or outsourcing them would be faster, and leaned toward outsourcing. So I sat an agent in front of my chats.
Using an agent in 2026 isn't just calling an LLM API. It's closer to placing an agent inside a runtime such as Claude Code. The real work is designing invocation timing, context boundaries, human intervention points, and runtime lifetime. Generating reply text follows from that.
This article records decisions behind auto-kakaotalk, an open-source Skill that learns my conversational style and replies on my behalf.
Choose the runtime first
The first question in an agent system isn't what it should do, but where it should run.
I considered three options: a Python daemon spawning claude -p, an OS scheduler using a launchd plist, and running directly inside a Claude Code session. The first two initially seemed natural because ‘daemon’ was familiar.
But this tool wasn't meant to be a 24/7 background service. It helps while I am not looking at KakaoTalk. A daemon would create an unattended period without a structural mechanism against unwanted messages. Logs and alerts would be needed, converting safety into operational work.
With the Claude Code session as the runtime, closing the session stops its work. It helps only while the window is open; the stop instruction is simply ‘close the window.’ I called this topological safety: a safety property derived from the system's arrangement and lifetime, not another feature to maintain. It limits when the system can act, rather than requiring another always-running guard.
That choice shaped the rest of the design.
Wake the agent with CronCreate
A session runtime needs a way to schedule work inside it. Claude Code's CronCreate is an in-memory scheduler valid for the session's lifetime. It accepts standard cron expressions and fires when the REPL is idle. During a conversation it waits; closing the session removes it.
Those properties fit this use case.
CronCreate("*/3 * * * *", "auto-kakaotalk tick — cycle.sh poll → draft → send")
That one line replaces what a traditional design would handle with a launchd plist, daemon process, health checks, and shutdown handlers. There is no separate daemon binary or restart script; changing the rule means editing one Markdown line.
Using it taught me that agent design is often more about exploiting a runtime than invoking code. Idle-triggered execution and session-scoped scheduling removed the need for a separate daemon architecture.
Keep the Skill thin
A Claude Code Skill begins with ~/.claude/skills/<name>/SKILL.md. You describe the task in natural language. Putting every procedure in one long document is tempting, but bad for context: the whole document loads every cycle, and longer instructions make attention less focused.
I kept a routing table and minimal loop in SKILL.md, moving details into references/*.md. They load only when the corresponding situation occurs.
| Task | Reference |
| Initial setup/failure | references/setup.md |
| Register a contact | references/register.md |
| Run the loop | references/loop.md |
| Persona overrides | references/persona.md |
| Two-phase sending | references/send.md |
Each cycle reads the table first and leaves irrelevant references unloaded. What the agent doesn't read matters as much as what it does. In this runtime, the context window acts as memory, making context allocation part of system design.
One place for judgment
A common agent pattern chains small classifiers: route → classify → choose a tool → generate. It suits some domains. This one called for the opposite.
I started with an emotion classifier. emotion_score.py examined exclamation marks, emoji, and length, returning a suspicion score. Above a threshold, the contact was considered angry and the reply model wasn't called. It broke within two days.
A friend sent ‘Fine.’ In context, it meant ‘Okay, do it your way.’ No exclamation marks, no emoji, short text: a low score, followed by a wrong reply. Meanwhile, ‘LOL that's insane hahaha’ scored high from repetition and was blocked, though it was just a joke. The conversation stopped there.
People don't run a separate classifier before reading a chat. One reading determines both the emotional context and whether and how to reply. Splitting those decisions tears the context apart. Script scores describe surface signals; the agent has the broader context.
I removed the classifier and kept one rule:
No script independently opens or closes the send gate.
Script signals are advisory inputs. The agent decides whether to send, with one LLM judgment per cycle. Splitting that integrated judgment undermined why I was using an agent.
Implementation kept tempting me to break the rule. In register.py, registering a contact gathers 500 historical messages. Automatically generating a persona while reading them sounded tidy. But it would create two judgment points: registration and the session loop, plus the question of whether those judgments agree.
So register.py doesn't call an LLM. It backfills the DB and creates an empty persona.md. The user and agent write the persona together within the session. Scripts gather data; the agent judges. Recognizing when that boundary was blurring became a central part of the project.
From N approvals to one
The difficult human–agent question is when to involve a person. Approval for every action makes the agent pointless; no approval at all can create trouble.
The first version used draft-and-ask. A cron tick found a new message, generated a draft, and asked in the session.
> [Maenggu] ‘I'm screwed’ (05:51 KST)
> Draft: ‘lol why’
> Send? (yes / no / edit: ...)
It looked safe. After a day, the problem was obvious. Typing ‘yes’ every time wasn't much less burdensome than replying myself. I'd replaced the one-line-reply burden with an approval burden. People can avoid the approval loop too.
I switched to automatic sending, concentrating approval into one calibration at registration.
During registration, the agent reviews 500 historical messages and writes a relationship report: tone, relationship dynamics, and topics it shouldn't reply to. Something like this:
[Maenggu relationship report]
▸ Relationship
- Familiarity: casual speech, jokes, shared self-deprecation; likely friends for 10+ years
- Initiation: 58:42 (Maenggu:me); Maenggu initiates slightly more
- Reply rhythm: me 12 minutes, Maenggu 4 minutes; I tend to be slower
▸ My style
- Average length: 1.4 sentences; short
- Frequent abbreviated endings and casual acknowledgments
▸ Recent changes
- Average delay rose to 38 minutes over the last 2 weeks. Something changed?
▸ Suggested no-reply topics
- Account details and transfers
- Long stories about dreams, which I historically didn't respond to
Is this accurate? What tone should I use when replying for you?
The user reads and corrects it: ‘I've been brief because of exams,’ ‘You missed that I call him an older brother,’ or ‘Don't write long sympathetic replies about dreams.’ The agreed corrections become state/personas/<chat_id>.md.
The loop no longer asks about every message. It uses the persona to decide and send, then leaves a one-line log.
> [Maenggu] ‘I'm screwed’ → Sent ‘lol why’
It's notification after the fact, not an approval request. Reducing N approvals to one was the most important decision. Autonomy versus safety wasn't just a binary choice; it depended on approval timing and frequency. Per-cycle approval and per-relationship agreement produce very different tools.
This matches a pattern I see in agent UX: tiny permissions asked repeatedly can make a tool unusable, while a few substantial agreements can support autonomy within understood boundaries. One substantial calibration asks the user to endure the setup once. Repeated micro-approvals demand effort forever. The nominal autonomy may be similar, but the experience isn't.
I didn't build the layers
I spent the most design time on a four-layer Reflection Stack, citing Reflexion, Voyager, Generative Agents, and MemGPT, with three layers of drift defense.
L1 Episode Raw conversation logs
L2 Reflection Daily summaries
L3 Lesson Reusable rules
L4 Persona Actual style guide
The document looked impressive: four papers, four layers, three drift defenses. During implementation, only L1 had a real use. What would L2 summarize, and for what purpose? L3 had no reflections to find patterns in. L4's persona merge approval had no lessons to merge.
They existed only in the future. They were plans, not features. Mistaking plans for features is common in agent design, and paper citations can make that mistake look well structured.
I removed them. What remained was conversation history in the DB and a persona file. Each cycle pulls the latest 30 messages and overlays the persona. The DB holds facts; the file holds intent. They meet only at judgment time. Four layers became two, without losing any actual functionality.
After deleting them, I stared at the diagram suspiciously. Could it really be this simple? Further use didn't reveal another required layer.
The DB provides the feedback loop
The immediate question was how learning would continue. A persona written once during calibration is static. What keeps the agent aligned as relationships and tone change?
The DB does. Each cycle calls db.py get-context --chat-id <id> --limit 30. Those messages form a sliding window, not a fixed sample. Today's jokes and yesterday's tone change inform the next draft. Sent replies also enter the DB and the next context window, so the agent reads its previous output as input to its next judgment.
This has two implications. Short-term changes can be followed without a separate learning pipeline, explaining why the Reflection Stack was unnecessary. And the persona's purpose becomes clear: specifying what intent should override observed behavior. Facts come from the DB; intent comes from the file; they combine at every decision.
Users guide long-term recalibration. If a formerly formal relationship becomes casual but persona.md remains formal, the user can ask to analyze the contact again. The agent reviews recent chats, the user corrects the report, and the persona updates. Automatic drift detection could be wrong and move the persona in an unwanted direction. Automatic short-term adaptation and manual long-term calibration define this system's improvement strategy.
Adapters as shell verbs
I isolated the platform boundary in one file: scripts/adapters/kakao.sh.
kakao.sh check → Check authentication
kakao.sh resolve → List chats
kakao.sh history <id> <limit> → Historical messages
kakao.sh poll <id> <since> → New messages
kakao.sh send <display_name> <text> → Send
Five verbs. I initially considered Python abstract classes for type safety and mypy. But the caller is an agent invoking bash kakao.sh poll ... and receiving JSON. A shared superclass offers it little benefit.
Stdout, exit codes, and JSON make shell verbs language-neutral. A Slack adapter could be Ruby and a Discord adapter Go. With an agent as caller, a cheap abstraction can go far. It also fits the session runtime: adapters return JSON, the agent judges, and invokes an adapter again. No deeper binding was needed.
KakaoTalk had no suitable entry point
The five-verb contract was clean. What sat underneath was not.
There was no documented official route for externally reading personal messages and freely replying in this scenario. Kakao's REST messaging APIs and business channels don't provide that personal-chat interface.
I used two unofficial paths: reading the local encrypted database and sending through macOS accessibility-driven UI automation.
The Mac app stores messages with SQLCipher. Its key is derived using PBKDF2 from user_id and the device UUID. The user ID isn't always stored plainly in a plist: some accounts have candidates, others only a hash. For the latter case, the existing authentication logic searches a range of up to a billion candidates in parallel, taking tens of seconds on an eight-core MacBook. This logic fits in the roughly 500-line _kakao_auth.py.
Sending presented a different problem. Without an official route, AppleScript operates the UI. Typing with a Korean IME active can produce the Korean character mapped to V rather than the intended shortcut. Instead, text goes onto the clipboard and is pasted with keycode 9 plus Command, bypassing text composition. Chat rows are located by name in the UI tree and double-clicked with cliclick dc:x,y; the input coordinate is fixed 70 pixels above the window bottom. These values came from experiments.
I didn't invent the whole stack. I reused the necessary parts of silver-flight-group's kakaocli and upstream authentication logic, inlining the IME-safe AppleScript. I kept only chats and messages, dropping unnecessary search and schema features. Choosing the thinnest sufficient implementation below the adapter was itself a design decision.
The agent sits on top of that stack and imitates social fluency.
Bias sending toward one kind of failure
UI-driven sending is irreversible. Once the automation presses Enter, it can't take the message back. Duplicate sending is the worst outcome.
‘Not sent’ and ‘sent twice’ have different costs. A duplicate apology makes things awkward. A missed message can be sent manually.
drafted --(phase 1: DB commit)---> sending
sending --(phase 2: adapter)-----> sent (normal)
sending --(phase 2 failure)------> failed
sending --(process crash)-------> (stuck)
(stuck) --(next-session recovery)-> failed
If the process dies between phases, the row remains sending. At the next session, cycle.sh check marks it failed. We can't know whether AppleScript pressed Enter successfully. If it did, the message appears as outbound on the next poll. If not, it stays failed rather than being automatically resent.
I'd call this safety-biased reconciliation. Decide explicitly which failure is less harmful and bias recovery toward it. That asymmetry belongs in a system sending messages on someone's behalf.
Unsolved problems
Concurrent Mac use is awkward: user typing can race with AppleScript. That's unresolved. Group-chat attribution also needs more real usage to determine whether the agent reliably distinguishes someone addressing me from one member addressing another.
Repeated feedback that silence was preferable could make the agent overfit to silence. A drifting baseline response rate is hard to identify. Automatic long-term drift detection is also unfinished: if the user doesn't notice a changing relationship, recalibration doesn't happen. The DB handles short-term feedback, but the system doesn't yet suggest when to recalibrate.
Conclusion
A few years ago, the question was how to prompt an LLM well. Now, it often looks more like how to arrange an agent system well. The frame has changed.
The LLM can generate reply text. My work was making the path from that text to an actual send thin and deliberate. I didn't change model capability. I chose the boundaries along five axes: runtime, lifetime, context, approval, and adapter contracts.
Project: auto-kakaotalk