Architecting a Scalable, LLM-Powered Attendance Management Platform | TeamSense

Oct 07, 2025

Architecting a Scalable, LLM-Powered Attendance Management Platform

At TeamSense, our journey to create a robust and scalable attendance management platform wasn't just about building software—it was about architecting a future. We envisioned a system capable of handling rapid growth and evolving user needs, with a solid foundation for innovation as crucial as speed.

Authors

James Stafford

Mitchell McKenna Principal Software Engineer

Fix the root cause of No-Call No-Show with help from TeamSense

Table of Contents

Strategic Beginnings and a Hybrid Proof of Concept Direct link to Strategic Beginnings and a Hybrid Proof of Concept

We began with strategic planning and an in-depth proof-of-concept (POC) phase. We explored various technologies, including Golang, Rust, and C#, each offering unique advantages. Our POC culminated in a hybrid approach: blending our existing monolith with serverless functions dedicated to Tier-1 services. This strategy allowed us to enhance, modernize, and strategically integrate LLM capabilities into our core functionality.

Our Core Platform Principles: A Blueprint for Scalability Direct link to Our Core Platform Principles: A Blueprint for Scalability

To guide our development and ensure long-term success, we established these core platform principles as the bedrock of our architecture:

  1. Generic yet Customizable: The platform provides a generic foundation that caters to diverse customers while enabling easy customization.
  2. Data as Platform: We prioritize data integrity, isolation, and accessibility, ensuring all data is discoverable, understandable, secure, trustworthy, and reliable.
  3. LLM Awareness: Our platform, services, workflows, and APIs integrate with LLMs to provide customers with intelligent recommendations and knowledge-driven capabilities.
  4. Data Contracts: Clear data contracts maintain consistency by defining structure, schema, data fidelity, and security requirements.
  5. Evolvable Designs: Our services and designs adapt to change, enabling updates without disruption while maintaining focus on ecosystem-wide solutions.
  6. Standards, Performance, and Security: We prioritize security through standardized APIs and logging protocols, while ensuring optimal performance through comprehensive Continuous Integration/Continuous Delivery (CI/CD) pipelines.
  7. Observability & Supportability: We build observability into our service design, with all production services maintaining logs, common identifiers, defined SLOs and SLIs, and metrics for requests, errors, and availability.

Architecting for Massive Growth Direct link to Architecting for Massive Growth

Our updated architecture is built for tremendous growth. To date, we've processed OVER 16 MILLION messages, conducted 1.5 MILLION surveys, and handled 400,000 messages EVERY SINGLE WEEK. With 21X revenue growth since 2021, our current investment strategy is focused on supporting and replicating this success at a hockey-stick growth trajectory.

Why Rust and Serverless? Direct link to Why Rust and Serverless?

Benchmark data provided by TechEmpower 2025-02-24 Single Query

After evaluating POCs in three languages, we selected Rust running on a serverless architecture. This choice was driven by:

Key Platform Components: Direct link to Key Platform Components:

Serverless Functions: Each endpoint, queue trigger, and processor operates as a distinct serverless function.

Databases: We use Postgres for relational data, Cosmos DB for document storage (such as conversations), and Redis for caching.

Job System: A robust job system built on Azure Queues handles our offline processes, including bulk messaging, with automatic retry capabilities (dead letter queue - DLQ).

Messaging Flow: Direct link to Messaging Flow:

Messages flow from our React front-end and SMS webhook calls through our existing Django app, which routes them to our new platform using message queues. When triggered, platform functions process these messages—performing database operations, integrating with external vendors (Twilio, Sinch, Mailgun), and handling logging and analytics.

Prioritizing Developer Productivity Direct link to Prioritizing Developer Productivity

Developer productivity lies at the heart of our platform. By streamlining development workflows, we accelerate innovation and enhance product quality. We've invested heavily in tools and processes that simplify local development, testing, and deployment.

Streamlined Local Development Direct link to Streamlined Local Development

To streamline local development, we've implemented several key commands and practices:

Collaboration Direct link to Collaboration

Infrastructure and CI/CD Direct link to Infrastructure and CI/CD

We utilize a robust CI/CD pipeline to automate our build, test, and deployment processes, ensuring rapid, reliable releases. Our infrastructure is managed as code, allowing us to create, modify, and destroy environments with ease and consistency.

Observability with Datadog Direct link to Observability with Datadog

We have implemented a comprehensive observability strategy using Datadog to gain deep insights into our platform’s performance, health, and behavior. Our goal is to detect and resolve issues proactively, ensuring a seamless user experience.

These observability practices allow us to maintain a reliable, high-performing platform and quickly address any issues that arise.

Conclusion Direct link to Conclusion

Our technical decisions have enabled us to build a scalable, secure, and adaptable attendance management platform. These architectural choices and principles will continue to guide our development, ensuring we deliver the best possible experience for our customers and are ready for future growth and innovation.