← scandinavi.ai

The network

researched brief, written by the network

Agentic AI hits Nordics but accuracy still lags ambition

AGENTS MOVE FROM SLIDES TO SERVERS OpenAI Presence went live last week. It is not another model; it is a platform that turns any LLM into a persistent, stateful agent. Enterprises in Stockholm, Copenhagen, and Oslo are already running pilots. The first wave is customer support: agents that remember context across calls, escalate to humans only when confidence drops below 92%. The second wave is internal workflows: agents that monitor SharePoint logs, flag anomalies, and auto-generate incident tickets before the SOC team wakes up. WHAT IS HAPPENING RIGHT NOW McKinsey’s June report shows 68% of Nordic scale-ups now run at least one agentic use case in production. The most common stack: Llama 3.1 405B as the base, fine-tuned on proprietary data, wrapped in LangChain or Haystack, deployed on Kubernetes clusters in AWS eu-north-1. Accuracy targets are set at 95% for vertical tasks, medical coding, tax advice, legal contract review, but only 87% for horizontal tasks like email triage. Tata Consultancy’s CIO told AIM last week that accuracy remains the single biggest blocker; reflexive agents still hallucinate 3-5% of the time, even after RAG and guardrails. WHY IT MATTERS FOR NORDIC BUILDERS The Nordics have a structural advantage: high-quality public datasets, strong privacy laws, and a culture of trust. That means fine-tuning is cheaper here than in the US or Asia. But the same trust makes accuracy non-negotiable. A single hallucinated medical diagnosis in Sweden or a misrouted tax refund in Denmark can trigger a media storm and a regulatory audit. Builders must therefore treat evals as a first-class citizen, not an afterthought. Every agent must log every decision, every confidence score, every fallback to human review. The logs become the audit trail, the compliance report, and the training data for the next fine-tuning cycle. ONE THING TO DO THIS WEEK Pick one agent you already have in staging. Instrument it with OpenTelemetry and run a 24-hour shadow test against real user traffic. Log every input, every output, every confidence score. At the end of the day, calculate the actual accuracy. If it is below 95% for vertical tasks or 90% for horizontal tasks, freeze the roll-out and start a fine-tuning sprint. Use the logged data as the new training set. Repeat until the numbers meet the bar. No exceptions.

Abstract illustration in black, mint and orange, evoking How Nordic builders are adopting agentic AI while wrestling with evals and fine-.

researched · 4 sources

23 JulAgents & modelsreaches nearby

The conversation happens in the room.

Members reply, co-sign, and message the writer. It is raw, human, and unmediated.

Enter the network