# Artem Frolov Source: https://artemf.dev/ Applied AI Engineer & Data Scientist Location: London, UK I build AI systems for real workflows. I'm an applied AI engineer and data scientist in London. I turn complex operational work into useful AI tools, from complaint investigation to pricing. My focus is reliable delivery: clear evaluation, traceable outputs, and people in control. ## Focus - Governed LLM applications - Agentic workflow design - Retrieval and context engineering - Databricks-native AI apps - LangGraph orchestration - MLflow evaluation and logging - MLOps for production decision systems - Internal tools for operational teams ## Currently Building AI assistants and decision tools at Domestic & General. ## Open to Applied AI, AI product, infrastructure, and agentic systems conversations. ## Links - [GitHub](https://github.com/artemf375) - [LinkedIn](https://www.linkedin.com/in/artemfr) - [Email](mailto:afrolov01@icloud.com) ## Impact - Live pilot Complaints Sidekick: AI investigation assistant moved from proof of concept into a Phase 2 live trial with complaint handlers (Source: complaints-sidekick) - ~80% Less manual analysis: Approximate reduction in manual analysis effort for internal pricing and analytics AI assistants (Source: experiences) - < 1 week ML deployment timeline: Model deployment reduced from approximately one month to under one week after MLOps improvements (Source: experiences) - ~4% First-year retention: Approximate improvement associated with a cancellation propensity model integrated into pricing decisions (Source: experiences) - Daily Transcript processing: Governed batch pipeline turns customer call transcripts into structured data for operational analysis (Source: call-transcript-analysis) ## Experience ### Senior Data Scientist — Domestic & General 2023-08 to Present · London, UK Building AI assistants, machine learning models, and pricing tools in a regulated insurance business. Leading work from business case and architecture through evaluation, governance, and rollout. - Led Complaints Sidekick from business case and architecture to a Phase 2 live trial, with complaint handlers reviewing AI summaries, evidence, and recommendations. - Built evaluation, monitoring, and feedback frameworks with usage logs, traceable outputs, quality assurance review, and human oversight. - Delivered AI assistants for pricing and analytics teams to investigate model performance and verify deployments, reducing manual analysis effort by approximately 80%. - Introduced CI/CD, automated testing, model versioning, and reproducible training pipelines, reducing deployment timelines from approximately one month to under one week. - Developed the organisation's first production cancellation propensity model for pricing decisions, with an approximately 4% improvement in first-year customer retention. - Delivered pricing optimisation using Earnix, simulations, elasticity modelling, and lifetime value analysis. - Worked with senior stakeholders across Pricing, Decision Science, Complaints, Operations, Risk, and Assurance to define requirements and deliver governed systems. Skills: Production AI, Agentic Workflows, Regulated AI, Data Science, Machine Learning, Pricing, Databricks, MLOps, LangGraph, MLflow ### Co-Founder / Builder — Paloma Labs 2025-11 to Present · Delaware, US Building analytics and machine learning tools for mobile apps, with a focus on event tracking, on-device prediction, and developer experience. - Built early Swift SDK prototypes for automatic and manual event tracking. - Designed event ingestion, session processing, and analytics schemas with FastAPI, PostgreSQL, and Redis. - Explored on-device inference for behaviour prediction and privacy-conscious product analytics. - Built web tools for configuration and analytics exploration. Skills: Swift, SwiftUI, FastAPI, PostgreSQL, Redis, React, Product Analytics, On-device ML ### Net Revenue Management - Commercial Excellence — Henkel Ltd 2022-08 to 2023-02 · UK Developed pricing analysis, retail reporting, and commercial decision tools for a consumer goods business. - Analysed Dunnhumby, IRI, and commercial data to identify pricing and revenue-growth opportunities. - Built reporting and decision tools, including electronic point-of-sale tracking and stock optimisation, that generated approximately £500k in commercial value. - Translated business questions into analysis of pricing, promotions, stock, and category performance. Skills: Revenue Management, Pricing Analytics, Power BI, Excel, Commercial Analytics ## Education ### Durham University — BEng, Electronic Engineering 2019-09 to 2022-06 · Durham, UK Studied electronics, communications, signal processing, control systems, mathematics, and applied statistics, with practical work in modelling and engineering design. - Led a six-person project on the engineering and commercial feasibility of torsion bar suspension for armoured vehicles; the project received a first-class mark. - Applied mathematical modelling and control theory to evaluate system performance and engineering trade-offs. Subjects: Statistics, Further Mathematics, Control Systems, Signal Processing, Electronics, Communications, Engineering Design ## Capabilities ### AI workflow discovery Work with operational teams to understand the task, map the current workflow, and decide where AI can make a useful contribution. - Discovery with Pricing, Complaints, and Operations - Business cases and feasibility assessments - Clear scope and success criteria ### Architecture & orchestration Connect retrieval, model calls, and structured outputs in a workflow with clear responsibilities for each step and the people who review it. - LangGraph orchestration for multi-step agent flows - Structured outputs and function calling - Next.js and Databricks apps ### Retrieval & context design Bring calls, notes, repair history, and operational data into a useful case record, with governed access and traceable sources. - Structured retrieval and summarisation workflows - Unity Catalog-governed data and transcripts - Batch and real-time context assembly ### Evaluation & observability Define quality criteria, log model runs, and review behaviour in use. Combine automated evaluation with user feedback and quality assurance. - MLflow evaluation and logging - LLM-based evaluation and user feedback - Usage logging, traceability, and QA review ### Governance & human oversight Give reviewers the evidence and controls they need. Build review steps and audit trails into the workflow with Risk and Assurance teams. - Human-in-the-loop controls and review-and-approve flows - Audit trails, usage logging, and traceability - Workflow design with Risk and Assurance ### Delivery & adoption Take a system from prototype to live use. Establish repeatable releases, deployment checks, and the stakeholder support needed for adoption. - CI/CD, model versioning, and reproducible training pipelines - Release governance and deployment verification - Stakeholder engagement and adoption ## Skills ### AI & Applied ML - Agentic AI Systems - LLM Workflow Design - LangGraph Orchestration - Retrieval-Augmented Generation (RAG) - AI Evaluation Frameworks - LLM-based Evaluation - Structured Outputs - Function Calling - Prompt Engineering - Feedback Loops - GBM / Propensity Modelling - Experiment Design & A/B Testing ### Business & Leadership - Enterprise AI Governance - Stakeholder Management - Executive Communication - Product Delivery - Regulated Workflow Design - Business Case Development ## Technology stack ### Languages & Data - Python - SQL - Spark - TypeScript - Swift - Power BI ### AI & ML - OpenAI API - Databricks - MLflow - LangGraph - LangChain - Scikit-Learn ### Web & Backend - FastAPI - PostgreSQL - React - Next.js - Tailwind CSS ### Infra & Tools - Docker - Nginx - Cloudflare - Tailscale - Git ## Projects ### Paloma ML analytics SDK An experimental Swift analytics SDK for app behaviour, with on-device prediction to explore the timing of prompts, offers, and paywalls. Technology: Swift, React, FastAPI, PostgreSQL ### Nightime Doodles iOS app SwiftUI app experiments with AI, document workflows, push notifications, and lightweight backend services. Technology: SwiftUI, Google Vertex AI, AWS Lambda, S3, APNs ### Raspberry Pi Homelab Self-hosting / infrastructure A Raspberry Pi setup for self-hosted services, private networking, reverse proxies, and deployment experiments. Technology: Docker, Nginx, Tailscale, Cloudflare, Hermes, OpenClaw ## Certifications ### Databricks Certified Generative AI Engineer Associate Databricks · 2026-06 Credential ID: 184791628 Verification: https://credentials.databricks.com/aac117df-d06a-43ab-940d-1a9b528e6c4a Certificate: https://artemf.dev/content/certs/nwxzwpua_1780992461967.png ### AI Agents Fundamentals Databricks Academy · 2025-10 Certificate: https://artemf.dev/content/certs/aieng_fundamentals.png ## How production AI ships From source data to reviewed, monitored decision support - Ingest: Calls, notes, history, operational data - Retrieve: Relevant context, governed access - Reason: Model workflows, structured outputs - Evaluate: MLflow evaluation, logs, feedback - Human review: Evidence, approval, audit trail - Deploy & monitor: Live rollout, monitoring, governance --- # Complaints Sidekick Source: https://artemf.dev/work/complaints-sidekick AI assistant / Insurance / Complaints Status: Production Designing, evaluating and piloting an AI assistant that prepares complaint evidence for human review. An operational pilot showed roughly 30% higher throughput, with quality monitored alongside productivity. An LLM can read a complaint history. The harder task is making that useful in a regulated business, where a missed detail can affect a customer’s outcome. ## The work before the decision A handler rarely starts with a clear account. Calls, case notes, repair records and earlier promises sit across several systems. Before making a decision, the handler has to reconstruct what happened. The Complaints Sidekick retrieved evidence, built a chronology and produced a concise briefing. It also worked as a chatbot in Teams, so handlers could ask for more information as their investigation developed. The handler retained responsibility for the investigation and the outcome. My role Led the business case, architecture, implementation and stakeholder engagement. Worked with Complaints, Assurance and senior stakeholders to move from proof of concept into a live trial. 01 / System design Evidence → judgement ### Support the investigation. Keep the decision human. **Complaint handler in Microsoft Teams** Entra ID · enterprise access Sidekick API **Explicit LangGraph workflow** Databricks 1. **Retrieve** Case records, transcripts, complaint guidance and SOPs 2. **Assemble** Organise relevant information and retain source references 3. **Respond** Prepare a briefing or answer a handler’s follow-up question MLflow traces retrieval, context assembly and model output Human review boundary **The handler investigates and decides** Validate facts · investigate gaps · assess the complaint · determine redress Simplified architecture. Handlers can request a briefing and ask follow-up questions in Teams. Responses give them pointers to source records and guidance; generated output cannot be used as evidence or determine the customer’s outcome. ## Build a briefing from evidence Long histories contain duplicate notes, conflicting accounts and uneven transcript quality. Sending all of that to a model does not guarantee a useful answer. I separated retrieval, context assembly and generation into explicit steps. LangGraph controlled what could be retrieved, when it could be retrieved and what reached the model. This made it possible to inspect each stage when a briefing was incomplete or wrong. The briefing needed to preserve source references, conflicting accounts and gaps in the record. A transcript records what someone said; it does not establish that the event happened as described. A pointer to the source > “The 12 March call transcript states that an engineer had visited earlier that day.” Illustrative wording. The handler still needs to find the transcript and check the account. This generated statement is not evidence. ## Continue the investigation through questions The briefing was a starting point. Handlers could ask follow-up questions and use the chatbot to find information in complaint guidance, documentation and standard operating procedures (SOPs). Retrieval-augmented generation (RAG) brought relevant material from those sources into the model’s context to inform its answer. This helped handlers find information during the investigation. They still needed to check the source guidance and decide how it applied to the complaint. ## Keep the handler able to challenge the output A recurring question during the work on call transcripts was how we could ensure that the AI was not making decisions. Giving it authority over complaint outcomes would have required further legal and governance work. Across briefings and follow-up answers, its role was to support the investigation. Responsibility for upholding or rejecting a complaint, determining redress and communicating the outcome remained with the handler. But even summarisation involves choices. The model selects what to include, what to leave out and what to emphasise. Those choices can direct a handler’s attention and affect their judgement before they reach a formal decision. Keeping the final decision human did not remove that influence. For this workflow, the generated output could not be used as evidence. It gave handlers pointers to relevant calls, events and records. They still had to find the source information, verify relevant facts and investigate gaps. The benefit was a faster route to that information, with the investigation and assessment still in human hands. Human review therefore meant more than accepting a provisional answer. Handlers needed to check the source material, consider what the summary might have omitted and challenge the model’s interpretation. This made omissions as important to evaluate as incorrect statements. Retrieved content also needed a clear trust boundary. A request in a transcript was complaint evidence, even if it looked like an instruction to the model. Constrained tool access and separate application instructions were controls against that risk; they were not proof that a model could never follow a malicious instruction. ## Check more than whether the answer sounds right These controls needed evidence behind them. I used several layers of evaluation to compare versions, find omissions and check whether the system helped handlers in practice. 01 Deterministic checks Response schema, required fields, retrieval success, unexpected tool calls, latency and execution failures. 02 Model-based review A separate evaluation prompt used an LLM to assess coverage, source consistency, unsupported claims and clarity against a rubric. Its scores identified cases for review and possible regressions. 03 Human assessment Human assessments provided a check on the LLM judge. Important changes still needed review against real examples; a model score alone did not establish correctness. 04 Operational outcomes The live pilot tested whether handlers could complete more work while maintaining acceptable quality and customer outcomes. ## Make feedback useful to engineering Teams let handlers rate an interaction and add a comment. That feedback was attached to the same MLflow trace as the response, so a poor rating could be investigated against the exact retrieval, context and model output. This made it possible to separate missing evidence from evidence the model had overlooked. Selected failures could then enter the evaluation dataset, helping later versions avoid the same mistakes. 02 / Learning from use Feedback → evidence ### A thumbs-down with a path to a fix Illustrative handler feedback > “Missed an important previous repair.” Feedback linked to the original MLflow trace Was the repair in the retrieved context? No / retrieval **Inspect the evidence path** Check the source data, API response and context assembly. Yes / generation **Inspect the model response** Check how the prompt and model handled the available evidence. 1. Save the case 2. Test a change 3. Run evaluations 4. Release & monitor A diagnostic example, not a recorded customer case. Selected traces become regression cases for later changes. ## Make the workflow ready for a live trial The technical controls were only part of moving beyond a prototype. I worked with Complaints, Assurance and senior stakeholders on data-protection and AI risk assessments, information-security review, production permissions, and monitoring and support ownership. Entra ID provided enterprise authentication and access controls. Configuration moved out of ad-hoc notebook state, deployment became repeatable, and tracing became part of normal operation. Prompt and model versions needed change controls so that updates could be evaluated before release. ## Test whether handlers could complete more work We ran a small operational pilot with a comparable control cohort. The analysis used individual baselines, trial-versus-control comparisons, several weeks of performance, case volumes, quality checks and handler feedback. Trial-cohort throughput moved from roughly 0.7 to 0.9 cases per unit of productive time: an increase of approximately 30%. Case complexity and handler experience vary, so a before-and-after chart alone cannot establish the effect of the tool. 03 / Operational pilot Rounded results ### More cases handled Trial cohort throughput **≈30% increase from baseline** Baseline **≈0.7** With Sidekick **≈0.9** 0 0.5 1.0 Cases per unit of productive time **Quality was monitored alongside throughput.** The available checks did not show an obvious material decline. Rounded trial-cohort values from a small operational pilot with a comparable control cohort. This chart shows the change from baseline, not the difference between cohorts. Some implementation details and figures have been simplified or rounded. ## Interpret the quality result with care Participating handlers already had strong quality scores. That left little room to demonstrate an improvement. The useful question was whether throughput could increase without signs of a material decline in quality. The result supported further use of the Sidekick to reduce preparation effort. It did not prove that quality was unchanged or that all customer outcomes were unaffected. The observed throughput gain was a promising pilot result, not a guarantee of the same gain at scale. The main lesson ## Remove effort before the point of judgement. The value came from giving handlers a better starting point. Explicit workflows, traceable feedback and layered evaluation made that assistance easier to inspect and improve. ## Tools & technology - Databricks - Microsoft Teams - Entra ID - LangGraph - MLflow - Genesys - Enterprise LLM endpoints --- # Customer Call Transcript Analysis Source: https://artemf.dev/work/call-transcript-analysis Applied AI / Operations Status: In operational use A daily Databricks workflow that turns customer call transcripts into structured data for teams to investigate recurring issues and sentiment. ## The problem Customer calls contain useful evidence about recurring issues, sentiment, and process failures. That evidence is hard to compare while it remains in long transcripts. Operations teams need structured data they can query and explore without reading each conversation. A simplified overview of the workflow. This diagram illustrates the approach; it is not a product screenshot. ## The approach - Stored transcripts in Databricks with access governed through Unity Catalog. - Built a daily batch pipeline to ingest and analyse transcripts. - Used structured LLM outputs to classify calls, assess sentiment, and extract operational signals. - Converted the results into SQL tables for analysis and reporting. - Built a Databricks app for Operations teams to explore trends and customer issues. ## What changed - Moved the workflow into regular operational use with daily transcript processing. - Made customer issues, sentiment, and process signals available as structured, queryable data. - Reduced reliance on manual transcript sampling when investigating customer conversations. ## Tools & technology - Databricks - Unity Catalog - Python - OpenAI Batch API - LLM structured extraction - SQL tables - Databricks Apps --- # The Age of One-User Software Source: https://artemf.dev/notes/the-age-of-one-user-software Published: 2026-09-05 AI agents are changing which applications are worth building. I built [Relay](https://github.com/artemf375/relay) with Codex because I wanted coding agents to reach me on my phone. It is a simple notification bridge, with a skill that agents can use to send progress updates and Live Activities. When an agent needs input from me, it can request it through the same service. Relay has a SwiftUI app and a small server that currently runs on my Raspberry Pi. It does a specific job for me. That was enough reason to build it. I also bought ESP32 devices for home automation and asked Codex to build me a home control panel. Within a couple of hours, I had three working devices. They now sit around my home, where I use them to control lights, scenes, and music. These projects are changing my first response to a software need. There may already be a suitable tool, but I increasingly want to build what I have in mind before researching what is available. I enjoy the process. For a personal project, that enjoyment is part of the value, even when building is not the quickest option. ## The minimum viable audience gets smaller Every custom tool has a cost: the time to build it, check it, and keep it working. The benefit has to justify that effort. When development takes less time, smaller problems can cross that threshold. A tool that saves a few people a recurring frustration may be enough. The minimum viable audience might be a family, a small interest group, or one person. They can judge whether the software is worth having by what it does for them, without needing thousands of other people to want the same thing. With one user, talking to myself now counts as user research. Personal software already exists. People have built it with spreadsheets, scripts, and small databases for decades. Robin Sloan described a related idea in [An app can be a home-cooked meal](https://www.robinsloan.com/notes/home-cooked-app/). What has changed for me is how much I can build with the time and tools I have. I see agents as a way to make that practice more accessible and extend its scope. More people may be able to turn a particular need into a usable application, although describing and checking the result still takes work. ## Choose the features that fit Codex has also built an app to help my parents watch movies, and it is working really well. They are not very comfortable with technology, and downloading movies for them takes some effort. The app uses my Raspberry Pi with SSD storage for the movies, and a small SwiftUI app showing what is available to download, what is already downloaded, and a way to watch it. Existing software has to serve people with different needs. A feature that I never touch might be the reason someone else chose it. But the settings and choices that support those needs can also make a simple task harder to find. A broader media app might already do everything my parents need. I built them an app that puts those few actions directly in front of them. Similarly, my home panels bring the controls I want together in the places where I use them. Fit can also mean combining things. I might want the download behaviour of one tool and the simple library view of another. A custom app could bring that combination together. It need not reproduce everything either tool offers, and it can use existing components where they already do the job. A small audience makes feedback direct. If my parents cannot find a downloaded movie, I can change that part of the app. The support team is also expected to attend family dinners. Fewer features could also mean less code and lower resource use, but neither follows automatically. Agents make extra features easier to add too. Keeping the app small remains a choice. ## Small apps will need somewhere to live I have a head start here: a Raspberry Pi that acts as my own server, with Cloudflare Tunnel and domains I control. Deploying another app is easy for me because that setup already exists. About a year ago, almost all my services ran on it, with some orchestration to manage them. Relay still runs there today. Not everyone has a server ready to use, or wants to learn how to run one. Some personal tools can live entirely on a phone or laptop. For those that need a server, getting an agent to build the app leaves another question: where does it go, and who keeps it running? The movie app needs more than somewhere to run. Its files need storage, access needs to be controlled, and the software needs updates. Those responsibilities add up as I build more tools. If I eventually have twenty small apps, how much time will I spend keeping them working? This raises an interesting question for cloud hosts. Much of the cloud conversation has been about helping an application scale to more users. What happens when an app only ever serves one family? The scale may move from users per application to the number of separate applications, each with its own dependencies, data, and maintenance needs. Many could sit idle for most of the day. My prediction is that this will create more demand for hosting with low idle costs and less manual administration. Small services can share infrastructure or run only when needed, but deployment, access control, backups, and updates also need to be easy to manage. A useful app for a handful of people should not require each person to become a server administrator. ## Someone still has to look after it Generating an app does not establish that it works. For the movie app, I need to check what happens when a download stops, storage fills up, or the server is unavailable. I also need to control access and keep the software updated. For tools that hold personal records, protecting data and checking that backups can be restored matter just as much as the interface. These responsibilities belong in the decision to build. Sometimes an existing, maintained tool will still be the better choice. Commercial and open-source tools remain useful options. But my own starting point is changing: I can try building the thing I want, and enjoy making it fit. Relay and the home panels already have a place in my daily life. An app that helps my parents watch a movie is now working well too. None of them needs a larger audience to be worth having.