The future of large language models (LLMs) has always been a topic of heated debate—open vs. closed; large vs. small; specialized vs. generalized. As these debates persist, aggressive fundraising continues. We like to think one step ahead, considering where and how LLMs could prove useful. We firmly believe that, for most use cases, LLMs will require humans in the loop. We will explore this further below.
Sometimes right, sometimes wrong
Source: Preplexity.ai
The example above highlights a key issue with LLMs: sometimes they are correct, and sometimes they are obviously wrong. Since LLMs rely on probability distributions to predict the next word or number, we don’t understand why certain conclusions are made. This makes it impossible to address the issue at a system-wide level. We can only correct errors for predefined use cases, while the rest must be manually checked. Without verifying accuracy in a world of endless possibilities, we will never have full confidence in the answers provided by LLMs.
Does 80% = 0%?
In theory, a car with three wheels is 75 percent complete. However, despite being close to 100 percent, we still can’t drive it. Therefore, from a user’s perspective, there is no practical difference between 75 percent and 0 percent. A similar logic applies to LLMs: if we need accurate answers, 80 percent or even 95 percent accuracy won’t suffice. Engineering and scientific use cases demand high levels of precision. When applying LLMs in high-stake situations, close monitoring or additional resources for quality assurance (QA) are essential.
On the other end of the spectrum, some tasks don’t require exact answers. For these, such as editing papers or writing letters, LLMs perform reasonably well.
Humans in the Loop
Although LLMs are used to automate processes and accelerate tasks, it remains essential to keep humans in the loop. Currently, this is the only way to ensure the accuracy of the output.
Tools that incorporate humans into the loop must visualize three components: workflows, decision points, and data feeds. After visually mapping the process, users could edit graphs leveraging visual editors and, once satisfied, approve the final version through an approval tool. Integrating Unified Modeling Language (UML) engines with AI would enable the visualization of architecture, design, and implementation of complex systems and workflows. For the system to function effectively, an LLM would need to convert tasks into workflows, while UML software would generate graphical representations of these workflows. This graphical representation would allow users to quickly review workflows, verify key decision points for accuracy, and make necessary changes before approving the execution of the workflow.
LLMs present an exciting opportunity to significantly increase productivity. However, just like fully robotic-operated factories with a "lights out" switch, realizing their full potential is still a long way off. LLMs continue to require humans in the loop.




