Designing and Operating a University Festival Platform
At a Korean university festival, concerts and student-run booths take over the campus. Departments and clubs also set up temporary pubs, serving food and drinks themselves. A student might spend the afternoon enjoying the festival and the evening managing tables at their department's pub.
Visitors need performance times and booth locations. Students running the pubs need to track available tables and call waiting guests. When the concert venue fills up, its organizers need to manage admission queues. These different kinds of work happen at the same time, across the same campus.
We built a platform connecting festival information with these on-site operations for Adelante, Kyung Hee University's festival at its Global Campus on September 28–30, 2026. Our team at the university's LIKELION chapter developed it in about a month. As engineering lead, I led system design, deployment, and production operations.
Connecting festival information with on-site operations
One website for performances and booths
Visitors could open a single website without installing an app. The home screen collected festival announcements, while the performance page showed each day's lineup and the current running order.
Notices and the campus map lived on the same website. Selecting a booth on the map opened its location and operating information, giving visitors one place to find what they needed on campus.
Table management and admission queues
For on-site operations, we provided a separate service called SPOT. Students running a pub could start, extend, and end a table's use from their phones, then call waiting guests when a table became available. The open-air concert venue's organizers managed registration and admission. The public website used these operational states to show pub and venue queue information.
Guests registered on iPads placed at the venues. We separated guest-facing registration from the screens used by student operators. Each pub could manage only its own tables and queue.
Following a queue from registration to the call
After registering, guests received a personal status link. They could check how many parties were ahead of them while exploring the festival, then return when a KakaoTalk notification arrived. Registration, status updates, and the call back to the venue were one product flow.
At the concert venue, a call led to admission instructions. Cancellation before a call and admission after it were separate transitions, so the guest's view matched the organizer's actions.
For pubs, the notification explained where to return and how much time remained. Deferring a turn and missing the call window also had their own states. This flow made table state changes and notification delivery part of the same design problem.
Photo verification for a campus event
In a sticker-hunt event, visitors found stickers around the festival and submitted photos. An AI model evaluated each photo; the festival API committed verification records and issued reward coupons. Visitors then used the coupons for a prize draw at the student council booth. Photo recognition and the physical reward were connected in one flow.
A month to build, three days to operate
I led the architecture, shared authentication foundation, pub operations, AWS infrastructure and deployment, and performance validation. I divided features into issues that gave teammates room to understand the full flow and make implementation decisions. I defined common design rules and service contracts, reviewed changes across the system, and brought those features through to production.
We had a month to launch, and both content and operational state would keep changing during the festival. Pub operations had to continue while we updated performance features. Calling a guest had to leave the queue state and notification work in agreement. The design started with three decisions: where changes should be isolated, which states must commit together, and how failures would be recovered.
Build only as much system as the problem needs. Deciding what “enough” means requires understanding both the work that must keep running and the operational cost the team can carry.
That principle shaped our choices. APIs that needed independent replacement were separated, while physical RDS and Redis resources were shared. We chose managed ECS rather than a cluster we would have to operate ourselves. We sized the database connection pool by measuring throughput. Every technical choice needed a clear problem to solve.
The result was a successful launch and three days of real use: 17,308 estimated unique visitors and 1,634,031 CloudFront requests. Thirteen pubs used SPOT, accumulating 1,386 hours and 45 minutes of table use across 648 sessions. That is time summed across tables, including concurrent use. The rest of this article explains the boundaries and failure-handling rules behind that operation.
Service boundaries followed data ownership
The visitor website brought information together. Operational permissions followed the responsibilities of the student council, pub operators, and concert organizers. Guest-facing iPads could register people but could not run the venue. Who could change which state was the starting point for service boundaries.
Separate the units that need independent deployment
A modular monolith would have simplified builds, deployment, and local integration. Our priority was the scope of changes and failures at each venue, rather than throughput. Replacing a pub server to update performance information would expose on-site work to an unrelated deployment. We wanted to control deployment, processes, and capacity separately for each domain.
We split festival information, pub operations, and concert admission into three Spring Boot API services. Each owned its contracts, domain model, migrations, and authentication. The visitor frontend composed their responses. This was a microservices architecture with independent deployment and explicit data ownership: services did not query each other's databases directly.
The split came with real costs: separate images, deployment pipelines, observability, and API contracts. We had to build that foundation within the same month as the product. Shared build and deployment rules in the monorepo, generated clients, and a common local launcher reduced repetitive work. We paid the integration cost before the festival so we could replace only the services that needed changing during it.
We also limited the split. Work that had to commit together stayed in one service, and we shared physical RDS and Redis resources at our initial scale. This kept change independence without multiplying storage infrastructure. Connection-pool limits, monitoring, and redundancy addressed the contention and shared failure risks that remained.
We used that boundary on the first festival day when adding cross-pub table inspection and control to the operations console. Deployment-selection validation chose only pub-waiting-api and god-console. It did not require replacing the other APIs or any of the six web apps.
Seating a party and starting its table session stayed in the same pub API transaction. Those two actions had to succeed together, so we kept them within one service.
An execution environment the team could operate
The team's Kubernetes experience made container orchestration a familiar approach. We needed task replacement, a maintained task count, and per-service resource controls. With a month to launch, AWS ECS gave us those capabilities within an operational scope we could manage.
The three APIs and operations console ran on Fargate. The AI service ran on ECS on EC2 because its model had different memory and startup characteristics. It was private, reached through Service Connect. The six React apps ran on S3 and CloudFront. We could adjust resources and deployments for five ECS services independently, with Terraform managing the infrastructure and normal versus festival capacity.
Redis held login sessions outside the tasks, preserving sessions through replacement without sticky routing. Each owning API checked permissions and pub scope. A cache failure could fall back to the database; a session-validation failure was fail-closed. Recovering a read and verifying authority required different failure policies.
The operations console's delegated JWT bound the HTTP method, path, body hash, and actor. The owning API atomically registered the jti in Redis to reject reuse. Centralized operational access still passed through each service's business checks.
Defining what concurrency and notifications must guarantee
A pub table serves several parties in one evening. An interrupted response can cause an operator's request to arrive again. Before choosing a mechanism, we defined which table session a command referred to and what a repeated command must preserve.
A lock cannot identify the session a command meant
In pre-launch concurrency validation, we found that a delayed request to end one party's session could end the next party's session. The lock was working: requests ran in order, and the late request changed the latest row. A lock orders changes. It does not establish which table session a request intended to change.
The operator screen sent the rowVersion it had displayed as expectedVersion on end and extension requests. Under the lock, the server compared it with the current version and rejected a mismatch with 409 STALE_TABLE_STATE. The command now referred to both the table number and the state the operator had actually seen.
The operator screen supplied the version. The operations console required it in a separate lockTarget check before invoking the same business method. The method permitted omission for those already-validated internal calls. However, the HTTP contract also left it optional, so an authenticated request without a version bypassed the stale-state check. Protection in the operator UI and console was not an API-wide guarantee. Closing that gap requires mandatory versions at the HTTP boundary and a separate path for validated internal calls.
Repeated call and seating commands preserved the committed state without creating another notification. The console used an operation ID and result ledger to identify retries. State versions reject stale intent; operation IDs identify repeated work.
Preserve the notification intent and track an uncertain result
Calling a guest includes getting the instructions to them. The database and an external messaging provider could not share a transaction. We therefore addressed three failure windows: a crash after saving the call, slow delivery submission, and a lost acceptance response.
We saved the call state and the final message content and sending intent in the same database transaction. A transactional outbox relay forwarded the work to SQS, and a Lambda worker submitted it to the provider. A process crash left the intent in the database, and the operator's request no longer waited for the provider. Workers claimed ledger entries and recorded results through the owning API.
Conditional claims handled duplicate SQS Standard deliveries. The harder case was not knowing whether the provider had accepted a submission. Retrying it could send the same instructions twice, so we separated the outcomes:
- A confirmed failure before submission allowed bounded retries, then the DLQ.
- Confirmed acceptance recorded the provider's receipt ID and prevented resubmission.
- An uncertain result after submission became
UNKNOWN. Automatic resubmission stopped while we reconciled the provider's history with the ledger.
Preserving the intent to send and preventing duplicate external sends are different guarantees. An explicit uncertain state keeps the evidence needed to make a recovery decision.
This policy accepted a possible notification delay while acceptance was checked, in exchange for avoiding blind duplicate submissions. Recording UNKNOWN was only the start of that process. The worker immediately reported the incident to the operations channel, and a five-minute scheduled job checked for unreported incidents and rechecked notification ledgers.
The console showed the current business state alongside the provider's history. Resending required a reason and acknowledgement of the duplicate risk. The owning API checked link validity and business deadlines, then created a new send ID while preserving the original record. A resend did not close the original investigation: the receipt ID or evidence of non-acceptance had to be recorded separately. We built the evidence trail and operational recovery path alongside the retry logic.
Across the festival, all 582 notifications obtained unique provider acceptance IDs on their first submission attempt. Final delivery to the recipient's device was tracked separately. Content translation also used source-version checks so a late translation could not overwrite newer content.
AI verification followed the same ownership rule. A teammate's YOLO and DINOv2 models evaluated photos; the festival API used participant locks and database constraints to commit sticker credit and coupons. We validated the first credit for a new participant too. Inference and reward commitment had separate responsibilities.
Deployment included a path back to a working release
We promoted validated main revisions to prod, selecting deployment targets from changed paths and shared-package dependencies. GitHub OIDC access was restricted to the repository and production environment. Application deployment permissions were separated from Terraform and business-data permissions.
Replace tasks while preserving business state
ECS prepared the new version while keeping healthy old tasks running. Login and business state stayed in PostgreSQL and Redis. Readiness and ALB target health governed replacement; a failed rollout could return to the previous task definition through the deployment circuit breaker.
We used that recovery path during preparation. A booth-image conversion change passed CI but exhausted the memory of two new API tasks. We restored the previous task definition, then removed server-side decoding and re-encoding that duplicated compression already done in the browser. The API retained bounded file validation. Deployment had to cover resource constraints and recovery from a failed release as well as functional checks.
During a festival API rollout on September 30, the running task count went 2 → 4 → 2, while one-minute observations showed at least two healthy ALB targets. We checked both completion of the deployment and the continued presence of targets able to serve requests.
Redundancy needs separate failure domains
A single point of failure does not disappear just because there are two tasks. If both run on one host, one host failure can take them down together. The VPC and ALB spanned two Availability Zones; the APIs and console had at least two tasks. AI tasks used distinctInstance and AZ spread across two EC2 hosts. RDS had a Multi-AZ standby, and Redis had a replica with automatic failover.
Shared storage still connected the services' failure exposure. We retained fail-closed session validation while providing recovery paths for those stores.
Alerts carried the evidence an AI agent needed
We also used Codex and Claude Code, with access to GitHub and AWS, in operations. An alert needed enough context to start an investigation and carry it through code inspection, changes, and recovery checks.
CloudWatch → SNS → Lambda enriched each Discord incident with logs from the relevant time and target, plus the host, error code, route, and request ID. An operator handed that incident to an agent, connecting evidence collection, investigation, changes, and recovery verification.
I handed over incidents, directed authorized work, and verified recovery. Code changes followed review and release procedures. Infrastructure changes required Terraform plan review and apply approval. We separated automatic evidence collection from agent execution, monitored the alert relay itself, and kept failed relay work in an encrypted queue.
Measure throughput before adding resources
During pre-launch load testing, database connection waits accumulated. A larger pool was an obvious candidate, but database CPU usage was low. We observed the JVM, container, and database together, then compared configurations by how many requests they actually completed.
With 0.5 vCPU, 1 GB of memory, four connections, and 40 virtual users, enabling virtual threads increased performance-query throughput from 16.6 to 45.5 RPS. Pub queries rose from 68.3 to 134.8, and concert-queue queries from 39.4 to 79.5. Platform threads on the performance path fell from 68 to 33, but connection waits barely changed, from 36 to 34. The throughput gain could not be explained by eliminating pool waits.
In a separate 64-user experiment, increasing the pool from four to 32 reduced waits but lowered throughput from 43.7 to 23.6 RPS. Database CPU stayed at 6–8%, usually with only one query executing. Involuntary context switches per request rose from 1.13 to 2.66. This supported increased application-side execution contention with the larger pool. We chose settings by throughput and latency without assigning an exact share of the slowdown to each cause.
We enabled virtual threads and kept the pool at four connections. Allowing more concurrent requests had to translate into more completed work. For AI, processing and queue limits bounded accepted work; readiness checks and pinned code and model hashes controlled startup and reproducibility.
Production gave us a baseline for the next festival
Real traffic peaked at 38.2 RPS averaged over one minute. That is measured incoming traffic, not a benchmark of maximum server capacity. Each of the three APIs and the console maintained at least two healthy ALB targets across 4,320 one-minute samples. No service-wide interruption was observed; 46 individual ALB or target 5xx responses were tracked separately.
All six primary read paths met the 500 ms server-processing p95 target during operation. The map measured 111.2 ms; the other paths were between 11.8 and 25.4 ms.
Across 14,617 browser LCP observations, p50 was 648 ms and p95 was 3,060 ms. LCP measures when the largest visible content appears. Its population and measurement boundary differ from server processing, so subtracting the two percentiles would not identify the cause. Meeting the server target still left a long tail in visible page loading. The next step is to relate LCP to device, connection, page, and image data.
CloudFront's weighted cache hit rate was 35.35%. We cached selected shared public responses, while map and booth details that needed immediate updates used no-store. Raising a cache hit rate also means deciding which information can tolerate a delay in freshness. We evaluated Redis session lookups separately from cache hits.
Before this first deployment, we had no production baseline. We now have business records, logs, and metrics to plan the next festival around actual demand by feature and time of day, and the state transitions people really used.
My role as lead was to bring the team's implementations under common business rules, failure-handling standards, and release procedures. I connected features across the system, aligning change boundaries, command validity, and the handling of uncertain sends. The team built and operated that design within a month, adding the boundaries we needed and keeping the operational scope within what we could carry.
A platform that worked in the field
We took a one-month build into three days of live use, supporting 17,308 estimated unique visitors and operations at 13 student-run pubs. Visitors used it to find performances and booths; student operators used SPOT to manage tables. The 1,386 hours and 45 minutes of cumulative table use put the system to work in real pub operations.
Under that demand, all six primary read paths met their response-time target, and we completed the festival with no observed service-wide interruption. All 582 notifications received provider acceptance on the first submission attempt. The service boundaries, business consistency rules, and deployment and recovery paths were implemented together and exercised in production. We launched as a team, supported real work on campus, and left a working system and measured baseline for the next festival.