Table of Contents
In this edition we want to share two ideas. First, Jev, the new AI model the industry has been talking about for the past two weeks, and why its value proposition is relevant for the market. Second, our answer to the question: will AI replace people? 🙂
We have been deploying agents for two years, and in our experience real business value from an AI agent still takes people and their judgement. Today we share seven decisions where human intervention remains key.
AI is evolving so fast that we are not sure this will hold in the near future. For now, every successful agent we have deployed has needed an experienced teammate making decisions along the way, arguing about them and getting them right after getting them wrong a few times first. Each round of iteration makes the agent better.
First, what Jev is.
Latest AI News & Trends

What is Jev
TypeSafe AI has just released a new AI model, Jev, the first of what it calls System One Models: a new class of AI model built for decisions inside software. Software sends Jev structured questions and gets back typed decisions, with probabilities and confidence it can act on.
TypeSafe's founder is Diogo Almeida, who co-invented Reinforcement Learning from Human Feedback (RLHF) at OpenAI, the method behind ChatGPT. According to the company, RLHF has produced LLMs optimised for human preferences, but it also causes problems such as overconfidence and unreliable answers. TypeSafe built Jev to be used directly by machines, with a new training algorithm called Reinforcement Learning for Calibrated Decisions (RLCD).
How Jev is different from an LLM
Every AI model most of us have used answers in text. Until now, getting the data we need out of an LLM has meant parsing the text it writes.
Jev is not an LLM. It does not write text, chat, write code or generate images, audio or video. It generates structured data. Jev understands language, but its output is always a decision, and every decision comes with a confidence estimate. Software can act on the high-confidence answers and send the rest to a human for review.
Most of what agents do inside a company is not writing. It's deciding. Route this ticket. Approve this step. Classify this document. That is why Jev is valuable.
The proposal of Jev is to replace certain uses of LLMs: classification, routing and automation. For those jobs it is faster and cheaper.
Jev accepts three types of question:
Noul: answers whether something is true or false, with a probability. For example: is this urgent? Does this support ticket need human intervention?
Choice: picks one option from a predefined list of up to 255 options, and returns the probability of each. For example: which team takes this ticket?
Score: places something on an ordered scale. For example: churn risk, low, medium or high.
Pricing is $0.042 per million input tokens, or $42 per billion, and output is free. In TypeSafe's own comparison, the models that would do this work today cost between $0.20 and $10 per million input tokens.
Jev responds in 70 to 500 milliseconds, against the 3 to 329 seconds a frontier model takes end to end, according to the benchmark TypeSafe cites. The company claims Jev runs 193x faster and 444x cheaper than LLMs on these tasks, based on its own evaluations.

Source: TypeSafe AI
What it could do for coworking operators
Some use cases for Jev:
Deciding whether a help-desk ticket needs escalating to the team or can be answered automatically.
Routing every incoming support request to the right team: billing, community or operations.
Rating each member's churn risk every week from plan type, bookings, check-ins and support history.
Scoring leads against the operator's own qualification criteria, such as role and location, then running a weekly automation that creates follow-up tasks for the highest-scoring leads.
Classifying every ticket into a predefined list of categories for analysis.
Powering faster browser agents. Jev Ultrafast, a browser agent built by Browser Use, completed a Zurich to London search on Google Flights in 7.1 seconds (screenshot below).

Source: Jev Ultrafast
Jev's speed opens a practical use case for operators: letting members find the best meeting room for a given hour and book it automatically.
Which other use cases could this unlock for a coworking space? We would love to hear them. 😉
Why AI Agents Still Need People: Seven Human Decisions Behind Every Useful Agent
On the question of whether AI will replace people, what we see in practice is that human decisions are key to generating real value from AI.
Every coworking operator has access to the same LLMs. The model is no longer the differentiator. Two coworking spaces can deploy the same support agent on the same day, for example, and get completely different results. What separates them is the quality of the decisions people make along the way, before launch and after it.
We have been deploying AI agents in production for the past two years. In every deployment, the part that made the agent work was the people behind it. That's why, while everybody is debating whether AI will replace people, we want to give that part the attention it deserves.
What follows are seven decisions where human judgement shaped the outcome most:
Decision 1: What the AI agent sounds like, its style and tone
Before an AI agent answers a single member, someone has to decide how it speaks. Brand voice covers the tone of communication and the language the agent should avoid. It also covers response style: how long an answer runs, whether it uses the member's first name, what structure it follows and how it closes.
Then there is the channel. Does email, chat or voice change how a reply is structured, or should every channel deliver the same experience? A chat reply can be concise, with a few steps to check. An email reply may need more detailed instructions, a different tone, or even access to tools that look up the member's account before answering. A voice agent cannot read out a numbered list at all. Whether the agent behaves the same on every channel or adapts to each one is a choice.
Who decides all of that? A human. 🙂
Decision 2: What the AI agent knows about the space
An agent is only as reliable as the content it draws on. For a coworking space, that content usually falls into these categories:
Space profile: type of space, member segments, plans, pricing, opening hours, locations, amenities, resources and meeting rooms.
Policies and rules: booking rules, guest access and cancellation terms.
FAQs: the questions the team answers most often.
Private services: custom products, pricing, and services offered to specific members.
Deciding which documents feed the knowledge base takes judgement. Exporting everything the operator has, every policy version and every outdated price list, performs worse. Someone has to decide what is current and what contradicts something else in the knowledge base. A well-curated list is the best place to start.
That decision does not end at launch. Prices change and a new location opens. Someone has to own the knowledge base, keep it in step with the space and check that every resource the agent draws on is still relevant. Otherwise the agent keeps quoting last year's rates with complete confidence.
Who decides all of that? A human. 🙂
Decision 3: What the AI agent can and can never do
Someone has to decide how the agent should behave in specific scenarios. What should it do when a member's question is vague? When should it ask a clarifying question to give a more precise answer? How should it answer on key topics?
Testing the agent across different scenarios before launch gives the team the input to write new guidelines, which help the agent produce more relevant answers.
Constraints are the things the agent must not do under any circumstances. In a coworking space, the list usually includes approving refunds or discounts, going off topic, promising service levels the team has not agreed, sharing access codes outside the approved flow and revealing anything about another member.
Who decides all of that? A human. 🙂
Decision 4: When the AI agent escalates
In which scenarios should the agent hand a ticket over to a teammate? The escalation strategy decides when the agent stops replying, who picks the conversation up and what the member is told in between.
The triggers vary by space. Some are topics: billing disputes, contract changes, a member mentioning they want to cancel their plan or naming a competitor. Others are signals, such as a member asking to speak to a human, giving negative feedback, asking the same question twice, or the agent's own uncertainty about its answer.
Who decides what a good handoff looks like? The person picking up the case should receive a usable summary: the issue, what the agent has already said and tried, and the member's account and location. Without it, they start from zero.
Who decides all of that? A human. 🙂
Decision 5: What a great answer and a failure look like
An AI agent learns what good looks like from a set of real, well-resolved cases. That set establishes the quality level the agent has to match or exceed, and becomes the reference for doing the job well.
Choosing which real past responses should become the standard depends on criteria that live in the heads of the team that has been answering members for years. Those criteria have rarely been written down, and nothing in the system can surface them. Someone has to sit with a list of resolved cases and mark the ones worth imitating.
Defining the structure of a successful answer gives the agent a step-by-step process to follow.
Weak examples matter as much as strong ones. Pairs of good and poor answers to the same question, with a note on what separates them, are what help the agent produce the most relevant answers.
Who decides all of that? A human. 🙂
Decision 6: When the AI agent is ready to go live
Without a clear definition of success, there is no way to evaluate whether the AI agent is performing well, and no way to decide when it is ready to go live for members. That calls for specific acceptance criteria, each with a simple score. Pass, average or fail is usually enough.
These are the nine criteria we use to evaluate the quality of the agent's answers:
Factual accuracy: every reply traces back to the operator's source content. No invented prices, limits, amenities or steps.
Completeness: the answer covers the full path to resolution, not the first half. If booking a meeting room for a guest takes four steps, all four are there.
Relevance: the answer addresses the question actually asked, not the neighbouring article the agent retrieved. This is the most common failure we see: high semantic similarity, wrong intent.
Clarity: readable on first pass. Numbered steps for procedures, screen labels named exactly as they appear in the member app, no unexplained internal jargon.
Tone and voice: matches the team's voice.
Correct action: the agent chose the right move from its available options: answer, ask a clarifying question, run an action or hand off. The choice is judged separately from how well it was executed.
Ambiguity handling: when a question has more than one plausible reading, or key context is missing, the agent asks one targeted question instead of guessing.
Handoff quality: the agent recognises what needs a human, escalates it and passes on a usable summary.
Boundary compliance: the agent does not commit the company to anything it cannot: refunds, discounts, contract changes or service levels the team has not agreed.
A great answer is not a property of the output. It is a position someone chose. Nothing in the model knows where that line sits for a given space.
The launch decision comes out of those scores. Some criteria are non-negotiable: a single invented price or a missed escalation is enough to hold the launch. Others can go live slightly below target and improve in production.
Who decides all of that? A human. 🙂
Decision 7: Which failures to fix first
Where does fixing the agent start, and which fixes come first?
We start by evaluating 50 real answers by hand. That manual review shows how the agent actually behaves and where it needs work, and it usually produces a long list of possible improvements. Automating it too early hides exactly what the team needs to see.
When the agent fails, and it will, the wrong output becomes the input for the next iteration. A well-documented error is the most valuable instruction available.
Most failures do not need a permanent rule. The ones that do are those that repeat across many replies. An operator who writes a new instruction for every imperfect answer ends up with an unmanageable list of fixes, and an agent that gets worse with each one. Someone has to separate the errors caused by missing content, where the knowledge base simply lacks the right document, from the ones that need a different tone or a new escalation rule, then decide which to fix first.
Order matters too. A wrong price quoted to a prospect costs more than a reply that misses the brand voice. The fixes that go first are the ones that cost the space money or a member's trust.
Who decides all of that? A human. 🙂
Final Thoughts
A successful AI agent is the result of many decisions and judgement calls that someone made, argued about and got right after getting them wrong a few times first. Continuous learning and iteration are a key part of improving it. That is why teams are fundamental to making AI initiatives relevant inside organisations.
The AI models will keep getting better. The decisions around them will still need someone who knows the space and its members.
That's it for today. See you in two weeks! 🙂


