<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>spectrabrain Blog</title>
        <link>https://spectrabrain.ai/en/blog_tech</link>
        <description>spectrabrain Blog</description>
        <lastBuildDate>Wed, 09 Sep 2026 19:00:00 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <item>
            <title><![CDATA[[POSTECH Industry–Academia Research] Vehicle Speed Estimation from Low-FPS CCTV]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/speed_estimation</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/speed_estimation</guid>
            <pubDate>Wed, 09 Sep 2026 19:00:00 GMT</pubDate>
            <description><![CDATA[[EVA × POSTECH Industry–Academia Collaboration] This research was conducted by Jaehun Hwang and Keonwoo Park of the POSTECH AIM Lab under the supervision of Professor Minseok Song.]]></description>
            <content:encoded><![CDATA[<p><strong>[EVA × POSTECH Industry–Academia Collaboration] This research was conducted by Jaehun Hwang and Keonwoo Park of the POSTECH AIM Lab under the supervision of Professor Minseok Song.</strong></p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="introduction-from-object-detection-to-motion-understanding">Introduction: From Object Detection to Motion Understanding<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#introduction-from-object-detection-to-motion-understanding" class="hash-link" aria-label="Direct link to Introduction: From Object Detection to Motion Understanding" title="Direct link to Introduction: From Object Detection to Motion Understanding">​</a></h2>
<p>Detecting a vehicle in a video is not enough to determine whether it is speeding or approaching a hazardous area. To estimate its speed in <code>km/h</code>, the system must connect observations of the same vehicle over time and convert its movement in the image into a real-world distance.</p>
<p>The goal of this research was to <strong>maintain stable vehicle identities in low-FPS CCTV footage, calibrate the camera view against real-world road coordinates, and generate speed information that EVA can use.</strong></p>
<!-- -->
<p>Rather than simply replacing a tracking model, the research team implemented and validated the entire pipeline—from associating vehicles at low frame rates to converting pixel coordinates into real-world distances and presenting speed results in a verification interface. The demo below shows vehicle IDs, trajectories, a measurement area, and estimated speeds together.</p>
<div class="div_center"><p><em></em></p><p align="center"><em>Vehicle speed estimation running in EVA</em></p><p></p></div>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-image-feature-extraction-describing-vehicle-appearance">1. Image Feature Extraction: Describing Vehicle Appearance<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#1-image-feature-extraction-describing-vehicle-appearance" class="hash-link" aria-label="Direct link to 1. Image Feature Extraction: Describing Vehicle Appearance" title="Direct link to 1. Image Feature Extraction: Describing Vehicle Appearance">​</a></h2>
<p>In low-FPS footage, a vehicle may travel far enough between frames that its previous and current bounding boxes no longer overlap. Because position alone is insufficient to identify the same vehicle, each detected vehicle crop is converted into a feature vector that can be compared over time.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="why-did-we-need-three-features">Why did we need three features?<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#why-did-we-need-three-features" class="hash-link" aria-label="Direct link to Why did we need three features?" title="Direct link to Why did we need three features?">​</a></h3>
<p>The visual cues that distinguish vehicles can be divided into shape, color, and contour. Each cue was reliable under different conditions, which made it difficult for a single feature to handle every CCTV environment consistently.</p>
<table><thead><tr><th>Feature</th><th>Strength</th><th>Limitation</th></tr></thead><tbody><tr><td><strong>DINOv2</strong></td><td>Captures vehicle type, body shape, and general visual appearance</td><td>Fine-grained differences may weaken in low-resolution footage or between similar vehicles</td></tr><tr><td><strong>HSV</strong></td><td>Quickly separates vehicles of the same type when their body colors differ</td><td>Sensitive to shadows, backlighting, nighttime illumination, and exposure changes</td></tr><tr><td><strong>HOG</strong></td><td>Uses contours and intensity gradients when color information is weak</td><td>Contours change under rotation, occlusion, and truncated bounding boxes</td></tr></tbody></table>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="dinov2-representing-overall-vehicle-appearance">DINOv2: Representing Overall Vehicle Appearance<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#dinov2-representing-overall-vehicle-appearance" class="hash-link" aria-label="Direct link to DINOv2: Representing Overall Vehicle Appearance" title="Direct link to DINOv2: Representing Overall Vehicle Appearance">​</a></h3>
<p>DINOv2 extracts a deep feature that combines shape, texture, and visual components. Early in the research, we also evaluated Qwen visual features because they aligned naturally with EVA's VLM direction. VLMs are strong at semantic similarity, however, while tracking must distinguish individual vehicles that may look almost identical. When we compared the score distributions for matching and non-matching vehicles, DINOv2 separated the two groups more clearly, so we selected it as the default deep feature model.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="hsv-quantifying-color-distribution">HSV: Quantifying Color Distribution<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#hsv-quantifying-color-distribution" class="hash-link" aria-label="Direct link to HSV: Quantifying Color Distribution" title="Direct link to HSV: Quantifying Color Distribution">​</a></h3>
<p>HSV separates hue, saturation, and value. After converting a vehicle crop from RGB to HSV, we accumulate H, S, and V values in predefined histogram bins to create a feature vector that represents which colors, and how much of each color, appear on the vehicle.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/hsv_color_space-862720611d1b5929d00cd894052f8eeb.jpeg" width="22%" alt="HSV color space"><img src="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAWMAAAFjCAYAAAGfxVp3AAAACXBIWXMAAC4jAAAuIwF4pT92AAAdrUlEQVR4nO3df2ycxZnA8XECsRNvbOOYxElD6jShDaCjAVciqhWJcpwIf1T3R+gfbVSJSE1712tOFIraFAloC4KoHKiHaK9Cqmnv+AtxOl11iqNKoUJnxK/owunED9mpsSNrDQnrbLqOYydpTrPJhrU9Gz9jz4xnXn8/EooVJuPXs88+3vedeWYUQrtwSc+BAxd8m+v3qAzJkhSj46rKFwd7esp/9vf3K3Xpa1/m+z3qKl9Uhl9f/F07dni9aP09+lf8XfnrHx5ec/nvS99/44r/rq6urny9aYdHCCO33FT+LqWOTWr3rmcvfsfOvZe/c+6Z2y5/faVRT3KkCQ+TSkjUsvvjZz/9P8JQITxC8RIetUJi7ZF3y39ua2hSXY//qPx100NPWvdPeITiLDxmyxJafuuN5T97269V6sNTc/5ei3OkZxvhyptPVY30fPFGDGVO4SF501W4ColqhEco4vBwnSXm8uu7gvAI5YrhMdunNVUVCnPNEt2rP/3gL70zz97DGn3XbKI/D1eUP/xYGmq+Rh0aGij/o8InWy7/485iTk2/HhPjw5rKg5Qp929VTl36AC9hyhL6gge/9HT56+qQqGYKj+w9rLn801fdHVeP+lxzbOVNp0Niv2GEZ3ssVkGeDsUYHpWXqfqBSXWozFUl7MpZonl0yveyQXiEcsVf49Uv3ZRQmaNKf/N9cE94hCK+CZjLu9yXJEcayJTZJtD7+vrmNPFuM2Ev+R4qxYxRzstXuv/SCoWCOqrvFS3ZTNRLv0f5/upC1boJE/2NN2/ebH3BtT5PPPdqx+WvKx9N155pUn37/nDli62rq0szJEKofriz+6GZM7HjJ85f/rSYqYcxhES10q9/qUb2PXDFNpVHDgPFUbW88+Hy15mapCckVHVGqHoYOeVRb9VzPdsnT4SEb8Em5BftFPG8Rni2UX1q5Li659KbbdGOcPZDwmbi/Qft16oux5PvhIRvCzLhvqgm27MVEpUwKLRcoyqTsrVe+vlOso8vaVOPXrqDztQ93RUfpFQm14v1DepYS2v5a5uJdT2JrgwT9PWXJtZV1eT6udMrVGfx4vjNOrFe/SCl+iFH9f3WRv3NLSfV3/7md9QdGzbO+HvTWgv9ICXfcHFJZa2QyMaDlClLCTo/fUN89y+PlL+2yqFVL32tFSzVJPMq5GHfrvj0svJAQ78h7uvqs76UwntbVOsN75e/ni0MJE9Ik3zTZetXc/XLtfX3u6w71+snDp+e/detDUIiCBdTUiaLftoLyKIQb3aXbK4t9M9RHR4ktwAY5ACCluMvtOrHGRWmuyLJOg4bRHIADHIAmUwXpjmPKTuGVDNUMlQvSprSxxzvronkAKZEcr9hlaFefWj6e5Ph4eGgF1/r2vR0zXR6gkE/s59OP2WebuJks1p7Zmb8ScdhuimDbHq2eNRyVeZcVnDO1TsP3q9yP398xr/OGfrTMziVSZFqlcf41QYujKqWrodn/H2th3mzpRHSRQDJ/OIzTuDXKOStntGtqN7sptp8JvOliOQAGOQAoksXNku8ajEt9Jjv5kfzQSQHwCAHsKDpYr6poXqRcDUfWzTNB5EcAIMcQJB04eITg4mPBeQ+EMkBTIlk09JU23ph/dBmhhq3vyb6l5nJDwzLcWstwTXRbeu/+Z2Z/6dqzWbF6MS4KpzbMuPvq/fHqjZbrfaUQTbujWKxZ4p+QUxPxWyYPi1oxrTQsUl1fXhU1nvHJuMSYxP9SPTU9e/P+D+VpXHTzTY+pIsA5vyLz/TLrLq2oZrpqZiq8Xk29Gfc6mXdFdX1E9WYfooYgxyAKF2YFoVM2XriksbiqGoyTPHkLWpgbNRvv1M17doj+he66MeUGmotblEN7p7YEckBMMgBiNKFsYrGtCikqoirmrd5NMONhKrxiaHWDqu1uNyAj0gOQLS4pfyLYBq9lGk6vShEr1kIRd/+Shes6MrWtWdmXrPJqsnGOS9kMREtbun7xsxlpKa1YnrVjWlRiC/6+YLp9td0I6FLh/PNMz8xmNLCXLdZq4V0EcCcb6trRcBc6hLnqrqecbZrC3GqXy1EcgAMcgDOp59CbvC8kCnABpEcAIMMAFicXujufkVaAR/iyPxqNlX7vnYOsFUdRHy6CIBBDoBBDoBBDuDys4uRkZFZF85V2CxCdMHmcBWbawv1c1we5Pb2dvnDlsAPZqxmKgSHwFQUrt6idr43c/rK9UMu0kUADHIADHIADHIAmdseZ/Ltt4wHMZkWSB4aG1D7L52QXa3W6a6sT44YgxwAgxwAgxzA5V98pVJJvMjOZoctF2x26Ro9e1YNGXbTajQsTNQLFsdPzFycaFpgqVzsppXL5cS3rrY7bLlg+n6mCqzGjk1qw8mZA2oqsxgonlQ7Deupl3fNXN+sauymlckDL1LEIAfAIAfAIAeQxG11Yc+9asTwy8zktdKYecscQy3hhN5hQFgHOB9EcgAMcgAMcgAMcgAMcgBp7NVpeBZRy5dzjepBNt9bfBjkABjkABjkABbsF5+LrSVNu3Td39auuubds1tEcgAMcgAMcgAMcgAMcgBJb1JtPOoi16jUCS/fbs6I5AAY5AAY5AAY5ACclpiVfv1L8z+Y5572e9asKx9aKFFrn3u9Q+10w8VRdciwFaXectLEtKe9ZMyclpiZVrjbMk3n/6JhqXF9m1GNfe6NWwAPKeM+94OrZ27op2rsaS8ZM9JFAAxyAAxyAAxyAE6PwKjF9HC91lEXpr8f8jgO0n3uFSVmcWOQA2CQA2CQA5jzbXXJ4lZZn34+Xa/h6LdaivUN87+tNtw+/6k4qorvuTsSrpZZb6tr7Xiye9c/iL+J6Xh5mzOkdV2esWzMcLSGfhZhulU2fYrQA7z/tOF2vdl8Cz/X3WpIFwEwyAEwyAEwyAHMelttusXUv31bb5j5i2T3xzN3R1EOzn6qdQSdSX5Zn+pePXM3llo/h+mXHFuWJYhBDoBBDoBBDmDWnVtMu5jkztUbdzwxHdvmQn6sJO7l9HidOivcjUX/HKa/d70rzaw7t/Ttm3lMnL6H3/meYceTTj/HxOkX9IM22QK3I2/erF4aXDrzfxgOll030Wz8+VwjXQTAIAfAIAfAIAfg9GTJWhuJzpf+BJA3/TIzuK9lmyp968eitnN9CG+LSA6AQQ6AQQ6AQQ6AQQ7AaYmZr0Nobc4ZCfWJwQaRHACDHACDHACDDAAAAExXZxqRRx955JmOjo6tktHSRRB6jb7rtnoViF6kkBKba/Y1bll/Pe7dvfsr4sYvdHe/stAH9/f19c3rIP+FYHPNvsYt669HrZjluQUygUBGJhDIyAQCGZmQuePtYd7RwaQwuEXtfGZmOYqJXlEbouRjrsjIyAQCGZlAICMTCGRkgvFmz8WW9vNtWygU1FHHBfq+2Vyzr3HTbfVWfRJ6L4DOs7JpZ72JQoyFCxXGQHZxbMB829pUhsTC6potxmLysZ+pLwqPYdAbekr3mzw0NqA+s33mNlgmOuhNG27UvA5PFUO18NECmUAgIxMIZGQCgYxMYIp6geiDCqXn6E1anMhgq9bey9PprSptdlK02XzKxY0hGRmZQCAjEwhkZAKBjEwgkJEJPLVwqLDnXjViceq+1GulMeMp/yb6HEnjMYcGE/oaTKf8J4iMjEwgkJEJBDIygUBGJnCzN4uRW26SN26RHYdr68u5RvXg1htF/6q3/Vqr8+uzgoyMTCCQkQkEMjKBQEYmEMjIhEX51MLqSYQnT40cF08739/WrroW/IrjRkZGJhDIyAQCGZlAICMTjDd7+iiqfuEeZnq/Mx9th4eHrcZXrwUW8zSV/HFupbjtnjXr1C8alora/t+qVWro3ISobbG+QQ0Jfz7dVldHS+THSuVts6T0xuBS0pi4EmMg6/PUpHuYHbXY78ymrWbTVryg3bMNwus41tIaRduNzfI39QdtJ8Rt84OyN6myfJ1r4aMFMoFARiYQyMgEAhmZEPUUdQxTyWuF08jaG7fHMZHc9NCTonb1QwPeryUUMjIygUBGJhDIyAQCGZlQZ/ohnnziiVe2bt16u+QHtDnJSLeVzjqNnj2rGv/3iHiM9bZSUroq2Ue/N61apZonzojaDjVfozYIp4dt215/c6eo7Z/05t3XymbrTo/XqSN1H4raap9p2SZuu339reK2O+6+2xizwY8nkx6zpdcLSINeky5S16Sl9ZrNcwi9HkJ8zR2bVNeHR720vWPDRuEFKzV4c7eo6dkTbeoli2nn0rd+LG7rAh8tkAkEMjKBQEYmEMjIBCdT1M+92iFqVxjcorYLp08bi6OqyWKtbF64uXUs6rffqZqEm2zrqWSbtt2r94raFj7ZovYfXiNqW14o3xDvnnJkZGQCgYxMIJCRCQQyMsF4szcyMqIO9vSIfj497Vy4eouora7CPTQmWwM7OjGuBoonxWM84fG8ZildlawLOiWG9ZTzkKzjP1m2Lb4nfz06z+ZEbXPn6tW6iWbZRSgljh9XnExR73xPNi2rB+0z258WtdWDvPMvj8h/zAiO2dKl9dKq5ENDA+Kp5PxIn3gqWQfx/tPy1+Nw8zFRW/3Uom/fH0RtFwIfLZAJBDIygUBGJhDIyAQnU9Sl778haqfvZHcelt9cLO+STbXGYnxJm1ou3FZKTw8Prn5f1PbtD29WLwmnkjuLOaWa5eu4pa+di/3ZfCIjIxMIZGQCgYxMIJCRCQQyMiH43m82d8lbf7/L+/W4pJ+0SDe41k8XDp+WPYm4r2WbuCr5oEVVe+j1ED6RkZEJBDIygUBGJhDIyISoN/qW3hjGwmYfvMV6U+YLGRmZQCAjEwhkZAKBjEwgkAEAAAAAAAAAQKYYT4x8obv7lYX+IUulksrlZNub6i1v9W6hC9nW5npTFMvr8aN9+74ianwhAn19feKL6DlwYMHb2lxvimJ5PUzxyjQ0kkcQI3kEMZJHECN5BDGSRxAjeQQxkkcQI3kEMZJnnHauNTMSUqFQUK2tslNAbTY28dXW5npTFMvr8b29e2fErHGXIPGJpR7Z/HDKYscdX22trtej517tEHf+Q+HhONqRr74YxethwscJJI8gRvIIYiSPIEbyCGIkjyBG8ghiJI8gRvIIYiTPOGMXwzkSeprzaH+/qK2eLas1mxOqrc31+lQY3CLuXR8kKfXWm2/G8XoYMO3sqK3PaeeRW24St93+0JPitq037BW3fbi3qPIfnRK1ffmGnzDtDNggiJE8ghjJI4iRPIIYySOIkTyCGMkjiJE8ghjJM87Y1ZreC2l4eFj83fSUr/SafbW1uV5bhZZrxP+isTgqbju+pE3cdtVko7itz9fDxBjEMVTtKovrOGox5eurrc31KsupZJv955ua5QH/3b88Iu/4r36ilredEDX9l0P/rvZ89AtR2/tW/q36+uZviNrWWrvBxwkkjyBG8ghiJI8gRvIIYiSPIEbyCGIkjyBG8ghiJI8gRvKM086YG5upZBtrj7wrbp1//Efitk0WldExIxMjeQQxkkcQI3kEMZJHECN5BDGSRxAjeQQxkkcQI3kEMZLHtPMspFPJuqzepirZRn7rjZ56zgYyMZJHECN5BDGSRxAjeQQxkkcQI3kEMZJHECN5BDGSRxAjeYtu2nny7bfUyL4HRG1LHZu8XYdVBTPTzldEJkbyCGIkjyBG8ghiJI8gRvIIYiSPIEbyCGIkjyBG8ghiJC8TB5QPv/Ab9eb+x0Rtj69Zp4aEB34X6xvEbT/OrRS1q3jj9i5x2yGrnuVsDjOfqGsWt82dq1drzzSJ2p5fMskB5do7E2fUhpOyF+RYS6uXtppN2xjYHGZev6QoPqC8dFWHyjecErVd2rjM6nB5Ez5OIHkEMZJHECN5BDGSRxAjeQQxkkcQI3kEMZJHECN5xhm7gz09C/5z6alkPRMnMaRnnoSVybZtn2qQTZ82Xb1CfbujVdTWp/rtd8p7HxoQN/1z3QY1frxN1FZPOXeq60RtJ8fHxPFWa3raGMR37dgh6tQnvRZCPI3bsUl1fXjUS9vXX39d1HT9+vWqq00W8D417drjpff3l/xOPO2cz+fU4eZjorbbV94qj7cawc7HCSSPIEbyCGIkjyBG8ghiJI8gRvIIYiSPIEbyCGIkL+i0s97gWkpXJetCTYne5Tn1Wlu7qO3KXJPVtPP9d94hanthRaPqjSAl1FtMJeeX9YnbflD8nDpXv07U9rZlXyjPxEmsKNalNe0s3aFd06Xy0mlnvb5BOj2sg7LrxIjsIjo2qa8L29pcr082087dq58Wt/1j7/Uqv1RWwfzyrd8Vx5AOYKadsegRxEgeQYzkEcRIHkGM5BHESB5BjOQRxEgeQYzkBZ12tjkrWW9wLZ12Xr90mdq2bZuo7YVr16jeXKOorU1ltM31+mQz7Vz4ZIu47efHVql1E7KNtsvTwzZTyQs97fzcqx3itrt3PStuO1AcVRulm0C/+LzqOiObEu0dL3mpjLaddm566ElxWxvdq/eKW+8/vEbcdm3jefHG2T/evNlqKplpZyx6BDGSRxAjeQQxkkcQI3kEMZJHECN5BDGSRxAjefOedi4MyqcuD43Jp0RHJ8bVQPGkqK3PTbZ9TTvbTA/bsJlK7izmxG31ec2ZnXbe+cwj4rb7vySvrr2976fyaWel1B0bNoraHRoa8NLWaprckq+pZNUsnyY/8tUXxWcw21YwM+2MRY8gRvIIYiSPIEbyCGIkjyBG8ghiJI8gRvIIYiTPOGNXa3rPRJ/jKzV+4ry4bX6sJG57cYpaNvvkq63N9doaXyI7U1lZvh42hoeHxa0LhYI4hmzbmhiDWDq9qPXt+4O4be6Z28RtH/3sb9QHwrOEC+e2qFPXv7+gbcdPtImv19ajNlPJwopkrfT9N8RtdaBJ4+Kox7YmfJxA8ghiJI8gRvIIYiSPIEbyCGIkjyBG8ghiJI8gRvIIYiTPOO3si8005/VP/I3KDy4VtdWl54dPy6ZmfbXVaxak12vLZtwWIzIxkkcQI3kEMZJHECN5BDGSRxAjeQQxkkcQI3kEMZJHECN5QaedbRz42q+8bOrsq61NNTDcIhMjeQQxkkcQI3kEMZJHECN5BDGSRxAjeQQxkkcQI3kEMQAAAAAAAAAAAAAAAAAAAABg0Xmhu/uVCx70HDjgo1tv/fb19ZX/8yG1sWCM6Xc6Xrup/dr8nqBgCQAiQDIGgAiQjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIAMkYACJAMgaACFzFiwDE7blXO5xfX2Fwi9r5zCPO+117pkkd+NqvnPe7GPDJGAAiQDIGgAiQjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIAMkYACJAMgaACFAODTgw+fZbamTfA86HstSxSe3e9azzfg+NDaj9X3raeb/jJ86ru1/6e5VvOOW875dv+InzPmPCJ2MAiADJGAAiQDIGgAiQjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIAMkYACIgLoculUqqv7/f+RUXCoWk+h0eHnbeZ0VqY5HiGA+/8Bv15v7HnPd7fM06NdRyjfN+i/UNaqA46rzf0YlxdXvfT533mx8rqXvXvKzqW8477/uf//tf1ff+55+c93vPii61afNm5/3q94cNcTLO5XJqs4cLPtrfn1S/FYxFmmP8zsQZteGk++R2rKXVW78bm90n+YHiSS/9avUtRbW87YTzfktXdXjZ82Jp4zJv7w8bPKYAgAiQjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIAMkYACJAMgaACIgr8EZGRtTBnh7nV1wup02o30qJo211jURqY+FzjMf/8z/K1XKuDemqs45NzvvtXZ5TTzU0Oe93/dJlSr34vPN+hzxV3+ky6z+f2qDGj7c573vtmSbVqa5z3u/k+Ji/3GZBnIzb29vVXTt2OL9g/WZOqd/KAHspA05sLHyOceH5X3opL9aJuOvDo867fa2tXb3++uvO+922bZvqOuO+BFiPwx0bNjrvVu+j8dE1Q17KofP5nDrcfMx5v9tX3urt/WGDxxQAEAGSMQBEgGQMABEgGQNABEjGABABkjEARIBkDAARIBkDQARIxgAQAcqhLVEOPbXfycd+5rzf0bNn1bk168qHcbrmqxx6Za5J3X/nHc77vXDtGtU7XnLerx6HQ0MDzvulHPpTlEN77pdy6Kn9fnHfA867rRx5n1I5dLnfEyPOu+3NNXq7XsqhL6IcGgBwGckYACJAMgaACJCMASACJGMAiADJGAAiQDIGgAiQjAEgAiRjAIgA5dCWKIee2m/JQ2lxsb6h/GdK5dAp9ks59EWUQ1dQDv0pT9es94/wUbZc8lRa7LMcWh+n7+MUZ70vxdc9lEP7LN+mHPoiyqEBAJeRjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIAMkYACJAMgaACIgr8EqlknV5n4QuL06p3+HhYed9Vvi6Zn3acqWqzSVdtuyj349zK533WdF09Qq1fv165/1eWNGY1BjrfnW1nGv5sZKaqGt23q+WO1dfLol27fySSW85yIY4GedyOS8lwHqPh5T6rfDR9zsP3q9yP3/ceb+NHZu8lBbrvSO8nOB8iY++v93Rqrra3L+he5f4uV5fY6z73djsPslr9S1FL+XQpas6VL7hlPN+lzYu85aDbPCYAgAiQDIGgAiQjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIAMkYACJAMgaACIjLofGpwp571YiPMmAPR7L7tvbIu86/g9474uXNn43tR10wTQ896fxb13s4ph/zwydjAIgAyRgAIkAyBoAIkIwBIAIkYwCIAMkYACJAMgaACJCMASACJGMAiADJGAAikOly6JFbbnLeZ8HDsekVT40cV/d4KC++v61ddTnv9aL81hud9+njaHrM1L16r/NRGV/Spn79wTUqP7jUed8vb/uWumvHDuf9Huzpcd7nXPDJGAAiQDIGgAiQjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIAMkYACJAMgaACERRDv3cqx3O+ywMblHbPZyq21i8eCp0U7P7kt0fvPi86vJQXtyba1TqhPNuk1S//U7VtGuP80vXpy376tdH2XLhky1q/+E1zvtde6bJeZ+LBZ+MASACJGMAiADJGAAiQDIGgAiQjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIQBTl0D/0UJbZWcyp1hv8nH579+g/Ou9XJVqq66NfnyXnamjAfZ/6lOxlfap79dPO+/VVtqzfH6p51Hm/2oGv/Upt3rzZeb+xnOLsC5+MASACJGMAiADJGAAiQDIGgAiQjAEgAiRjAIgAyRgAIkAyBoAIkIwBIAIkYwCIgLgculQqqf7+fudXXCgU1JGvvui837fefFM93Ft03u+qyUY1cd1vVf0S933/uW6Den/J7xZ9vxN1zeU/UxrjD4qfU3/svd55v58fW6XWNp533m/uXL2X993w8HD5Px90rvCVg3z1a0OcjHO5nJd686P9/d76zX90ynm/Wn1LUS1vc3/2/fjxNvqtktJYnKtfp/JL3cfbuolmlW/w06+P911FarnCV782eEwBABEgGQNABEjGABABkjEARIBkDAARIBkDQARIxgAQAZIxAESAZAwAEaiTXsKTTzzxytatW293fcn9nqpffPWrSxz/7d3/UqWrJpz3vfZMk5dqq9T61aW6WkpjfNuyL6i/vnW7835TfH9ora2tzvtObSx0v9/bu1ecY8Xl0O3t7equHTvmfGE19fQk1a8e4IcHu728oTvVdepw87FF369OmFpKY7x95a28Py69P5SncujUxkL3a4PHFAAQAZIxAESAZAwAESAZA0AESMYAEAGSMQBEgGQMABEgGQNABEjGABABcQXeyMiIOmhZUSJRrthJqF9d7nnPii61tHGZ874nx8fKlVyLvd/zSybLf6Y0xiuKdbw/qsqhbQ/jlEhtLHycOF32Qnf3Kxc86DlwwEe33vrt6+sr/+dDamPBGNPvdLx2U/u1ybE8pgCACJCMASACJGMAiADJGAAiQDIGgAiQjAEgAiRjAIgAyRgAIkAyBoAIUA5tiXJP//0yxun2y2s3rV8AAAAAAAAAAAAACVJK/T+v1AqFMWeY4wAAAABJRU5ErkJggg==" width="17%" alt="HSV histogram binning example"></div>
<p><em></em></p><p align="center"><em>The HSV color space (left) and an example of accumulating pixel values by histogram bin (right)</em></p><p></p>
<p>We compare two vehicles' color distributions with histogram intersection, which sums the smaller value in each corresponding bin. Greater overlap produces a higher score, helping separate vehicles with similar types but different body colors.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="hog-quantifying-contours-and-intensity-boundaries">HOG: Quantifying Contours and Intensity Boundaries<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#hog-quantifying-contours-and-intensity-boundaries" class="hash-link" aria-label="Direct link to HOG: Quantifying Contours and Intensity Boundaries" title="Direct link to HOG: Quantifying Contours and Intensity Boundaries">​</a></h3>
<p>HOG converts each vehicle crop to grayscale, then calculates the direction and magnitude of brightness changes at every pixel. It builds a histogram of gradient directions within small cells, normalizes neighboring cells together as blocks, and concatenates the results into a contour feature that relies less on color.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/hog_gradient_histogram-4f607b565fc1cf4900967e86dc423b25.png" width="70%" alt="HOG gradient histogram construction"></div>
<p><em></em></p><p align="center"><em>HOG represents the direction and strength of brightness changes as a directional histogram<br>Source: <a href="https://learnopencv.com/histogram-of-oriented-gradients/" target="_blank" rel="noreferrer">LearnOpenCV</a></em></p><p></p>
<p>No single feature was sufficient for every scene. HSV was especially effective when daytime color information was clear, while HOG contours provided a useful fallback when color cues weakened. We therefore normalized and fused the three scores using a baseline ratio of <code>DINOv2 0.3 : HSV 0.5 : HOG 0.2</code>.</p>
<p>This fusion allows the remaining features to support a match when one signal is disrupted by lighting or occlusion. Although the ideal weights can vary by scene, using all three provided a stable common baseline across different CCTV environments. The fused representation acts as a visual fingerprint for comparing past tracks with current detections.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-low-fps-tracking-linking-ids-with-appearance-and-motion">2. Low-FPS Tracking: Linking IDs with Appearance and Motion<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#2-low-fps-tracking-linking-ids-with-appearance-and-motion" class="hash-link" aria-label="Direct link to 2. Low-FPS Tracking: Linking IDs with Appearance and Motion" title="Direct link to 2. Low-FPS Tracking: Linking IDs with Appearance and Motion">​</a></h2>
<p>Connecting vehicles based only on similar image features can swap their IDs when vehicles of the same color and type appear together. Motion alone also fails when a vehicle has just entered the scene or when the interval between frames is long. We therefore designed the tracker to <strong>use appearance as its primary matching signal and motion as supporting evidence.</strong></p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/feature_tracking_flow_en-6ee6c7cda103745ab3120bd554572132.png" width="95%" alt="Feature-based low-FPS tracking flow"></div>
<p><em></em></p><p align="center"><em>Feature similarities are fused, and only candidates that pass motion and score gates are linked one-to-one</em></p><p></p>
<p>First, the tracker calculates fused-feature similarity for every combination of past track and current detection. For tracks with enough history, it adds a motion prior based on recent direction and speed to check whether the vehicle could plausibly have reached the observed position. Appearance finds candidates across large movements; motion checks whether each candidate is physically plausible.</p>
<p>Hungarian matching then assigns candidates globally without duplicates. An assignment is not accepted automatically. The tracker checks whether the similarity is high enough, whether it is clearly better than the next candidate, and whether the track and detection rank each other first.</p>
<p>If a pair fails those checks, the tracker removes that combination and searches the remaining candidates again. This recovers valid alternatives hidden by the first assignment without forcing an uncertain link. Only accepted observations update track memory, while new detections and tracks that have left the scene are managed separately to keep a single false detection from contaminating later matches.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="improving-tracking-performance-at-night">Improving Tracking Performance at Night<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#improving-tracking-performance-at-night" class="hash-link" aria-label="Direct link to Improving Tracking Performance at Night" title="Direct link to Improving Tracking Performance at Night">​</a></h3>
<p>At night, both color and position cues become less reliable. Emergency lights and headlights can sharply change the HSV distribution of the same vehicle from one frame to the next. A detection box may also follow glare or a road reflection instead of the vehicle body, making the vehicle appear to move against its true direction. Because the next observation is already far away at low frame rates, neither appearance nor motion alone can reliably recover from this error.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/night_glare_bbox_distortion_clear_bbox-001b59f1380e121bf13a26a11696ca0b.png" width="80%" alt="Nighttime tracking sequence with glare-distorted bounding boxes"></div>
<p><em></em></p><p align="center"><em>Frames shown in chronological order. As the bounding box follows a glare blob, the observed ground-contact point moves opposite to the vehicle's direction of travel</em></p><p></p>
<p>Rather than relaxing every threshold, we adjusted the feature combination, candidate comparison, and track-retention policy for the causes of nighttime failure.</p>
<table><thead><tr><th>Change</th><th>Nighttime configuration</th><th>Intended effect</th></tr></thead><tbody><tr><td>Candidate competition</td><td>Excludes candidates already assigned to another pair from margin and mutual-best competition</td><td>Prevents a correct match from being rejected because of a candidate that can no longer be assigned</td></tr><tr><td>Fusion</td><td>Removes HSV from weights, thresholds, and absolute floors; uses <code>DINOv2 0.8 : HOG 0.2</code></td><td>Limits the effect of abrupt HSV changes while glare appears or clears</td></tr><tr><td>Track lifecycle</td><td>Reduces lost-track retention from five frames to three</td><td>Prevents stale tracks from continuing to distort normalization and competition scores</td></tr><tr><td>New tracks</td><td>Lowers the detection-confidence threshold from 0.5 to 0.4, but issues a new ID only when track confidence is at least 0.6</td><td>Reduces nighttime missed detections while suppressing unnecessary IDs from false positives and partial detections</td></tr></tbody></table>
<p>At 1 FPS on 12 UA-DETRAC sequences, these changes raised IDF1 from <strong>64.0 to 75.3</strong> and MOTA from <strong>56.5 to 64.8</strong>, while reducing ID switches from <strong>640 to 176</strong>. Replacing the deep feature accounted for about 15% of the improvement; the remainder came from the revised competition rules and nighttime-specific gates and lifecycle. The result reinforced that good features alone are not enough: <strong>the tracker also needs a consistent policy for when to reject and when to revisit uncertain associations.</strong></p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-camera-calibration-converting-pixels-into-meters">3. Camera Calibration: Converting Pixels into Meters<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#3-camera-calibration-converting-pixels-into-meters" class="hash-link" aria-label="Direct link to 3. Camera Calibration: Converting Pixels into Meters" title="Direct link to 3. Camera Calibration: Converting Pixels into Meters">​</a></h2>
<p>Tracking provides the pixel position of a vehicle and the time of each observation. Perspective, however, means that the same pixel displacement represents different real-world distances near and far from the camera. Pixel coordinates must therefore be projected onto a road plane measured in meters before speed can be calculated.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/pixel_distance_perspective-c50d6317355d9b1fec648e66c2188be7.png" width="50%"></div>
<p><em></em></p><p align="center"><em>Perspective example: the same real-world distance of 5 m appears as different pixel lengths</em></p><p></p>
<p>The speed-estimation flow is straightforward.<br>
Tracking first connects positions and capture times for the same vehicle, while calibration projects the vehicle's position from the image onto the road plane in meters. Finally, the system compares the real-world distance traveled with the elapsed time to estimate speed.</p>
<p>We initially investigated automatic calibration using monocular depth, an assumed average vehicle length, and vanishing points. Although convenient, this approach was sensitive to small estimation errors in CCTV scenes combining low frame rates, turning vehicles, and indistinct lane markings. It could also produce plausible but incorrect results whose cause was difficult for an operator to verify.</p>
<p>The final design therefore uses <strong>user-defined calibration</strong>. The user identifies reliable geometric references on the road plane, while the system calculates, validates, and stores the homography. It supports two input methods:</p>
<p>A homography unfolds the road area, which appears trapezoidal in the image, into a top-down plane. By matching image points with known road-plane coordinates, the system can convert vehicle positions from pixels into meters.</p>
<ul>
<li>Four corners and the real length and width of an area that is rectangular on the road plane</li>
<li>At least four reference points with known real-world <code>(X, Y)</code> coordinates, including points matched to a site plan</li>
</ul>
<p>BrnoCompSpeed is a public benchmark dataset for vehicle-speed estimation that provides fixed-road-camera footage with ground-truth vehicle speeds.<br>
Using actual tracking trajectories from nine of its camera views, we measured the following mean absolute errors.</p>
<table><thead><tr><th>Calibration reference</th><th style="text-align:right">Mean absolute error</th></tr></thead><tbody><tr><td>Dataset ground-truth area</td><td style="text-align:right"><strong>1.13 km/h</strong></td></tr><tr><td>Short rectangle</td><td style="text-align:right">7.48 km/h</td></tr><tr><td>Long rectangle aligned with travel direction</td><td style="text-align:right"><strong>1.52 km/h</strong></td></tr></tbody></table>
<p>The resulting guideline is straightforward.<br>
Place the calibration area where vehicles actually travel and make it as long as practical in the direction of travel. A one-meter input error creates a much larger relative error over a short area than over a long one. For projecting each vehicle onto the road plane, a single point at the <strong>bottom center of the bounding box</strong> was more stable than averaging multiple points.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-estimating-real-world-road-dimensions-with-a-vlm">4. Estimating Real-World Road Dimensions with a VLM<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#4-estimating-real-world-road-dimensions-with-a-vlm" class="hash-link" aria-label="Direct link to 4. Estimating Real-World Road Dimensions with a VLM" title="Direct link to 4. Estimating Real-World Road Dimensions with a VLM">​</a></h2>
<p>User-defined calibration requires the real length and width of the selected road area. Direct measurement remains the safest source, but an operator may not know the appropriate region or initial values when configuring a new camera. We therefore evaluated whether a VLM could inspect the marked road area and surrounding scene, then suggest initial real-world dimensions.</p>
<p>In the experiment, the four corners were labeled <code>P1–P4</code>, and the measurement area was emphasized with a blue outline and semi-transparent red fill. The most stable results came from providing <strong>both the annotated image and the four pixel coordinates while keeping the question concise</strong>. An image alone or coordinates alone performed less consistently.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/vlm_overlay-973e7893b1edd8dbfa18327407e8aec8.png" width="85%"></div>
<p><em></em></p><p align="center"><em>Example road-area overlay provided to the VLM</em></p><p></p>
<p>Using <code>Qwen3.8-27B-FP8</code>, we evaluated the improved input format on 21 images spanning three camera views across seven scenes.</p>
<p>Mean ratio error measures how far the VLM's estimated length and width deviate from the actual values on average; lower is better.</p>
<table><thead><tr><th>Evaluation metric</th><th style="text-align:right">Initial method</th><th style="text-align:right">Final input method</th></tr></thead><tbody><tr><td>Mean ratio error</td><td style="text-align:right">50.16%</td><td style="text-align:right"><strong>25.84%</strong></td></tr><tr><td>Mean absolute length error</td><td style="text-align:right">-</td><td style="text-align:right">7.80 m</td></tr><tr><td>Mean absolute width error</td><td style="text-align:right">-</td><td style="text-align:right">1.83 m</td></tr><tr><td>Both dimensions within 10 m</td><td style="text-align:right">-</td><td style="text-align:right">18/21 (85.7%)</td></tr></tbody></table>
<p>The revised input nearly halved the mean ratio error, but a 7.80-meter mean length error was still too large to use as the definitive metric scale for speed estimation. Longer perspective instructions and explicit reasoning did not improve every scene consistently.<br>
At this stage, the VLM is therefore better suited as a <strong>configuration guide that a person can verify</strong> than as an automatic source of ground truth.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-speed-estimation-using-capture-timestamps">5. Speed Estimation Using Capture Timestamps<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#5-speed-estimation-using-capture-timestamps" class="hash-link" aria-label="Direct link to 5. Speed Estimation Using Capture Timestamps" title="Direct link to 5. Speed Estimation Using Capture Timestamps">​</a></h2>
<p>The final speed estimator uses three pieces of information together:</p>
<ul>
<li>The track ID links observations of the same vehicle.</li>
<li>The calibration result converts the bottom center of the bounding box into metric road coordinates.</li>
<li>The actual capture time provides the elapsed time instead of relying on frame count.</li>
</ul>
<p>The estimator maintains a short position history for each track, then checks how far apart the first and last points in the recent window are on the road plane. It calculates average speed by comparing that real-world distance with the actual capture-time difference.<br>
Using capture timestamps preserves the correct time basis even when the input frame rate changes. An exponential moving average reduces small fluctuations caused by bounding-box jitter.</p>
<p>The current output is not an instantaneous speed. It is the <strong>average ground-plane speed</strong> between the endpoints of the recent window. Curved paths may therefore be underestimated, while camera shake and non-planar roads require additional handling.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/demo-9a0b73b522e97d904d8e66a084fb7595.png" width="100%"></div>
<p><em></em></p><p align="center"><em>Vehicle speed estimation demo</em></p><p></p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="6-productization-roadmap-from-research-to-an-eva-feature">6. Productization Roadmap: From Research to an EVA Feature<a href="https://spectrabrain.ai/en/blog_tech/Research/speed_estimation#6-productization-roadmap-from-research-to-an-eva-feature" class="hash-link" aria-label="Direct link to 6. Productization Roadmap: From Research to an EVA Feature" title="Direct link to 6. Productization Roadmap: From Research to an EVA Feature">​</a></h2>
<p>The purpose of this research was not limited to calculating vehicle speed. It was also intended to establish the foundation of an EVA feature that site operators can configure and use in practice. Combining vehicle speed and direction with EVA's detection zones, object information, and event rules enables the system to recognize scenarios such as the following.</p>
<table><thead><tr><th>Scenario</th><th>Example application</th></tr></thead><tbody><tr><td><strong>Speeding detection</strong></td><td>Alert when a vehicle exceeds the limit on factory roads, logistics sites, or campuses</td></tr><tr><td><strong>Hazardous approach detection</strong></td><td>Detect a vehicle rapidly approaching a worker, facility, or restricted area</td></tr><tr><td><strong>Sudden acceleration or deceleration</strong></td><td>Identify abnormal driving behavior involving a sharp speed change</td></tr><tr><td><strong>Abnormally slow movement or congestion</strong></td><td>Detect vehicles moving too slowly or remaining congested in a travel zone</td></tr><tr><td><strong>Operational analytics</strong></td><td>Show speed distributions and hazardous-event trends by time and zone in a dashboard</td></tr></tbody></table>
<p>This research was not limited to validating technical performance. <strong>It was conducted with productization in mind, considering the UX through which users configure the feature and verify its results.</strong> An operator can draw the vehicle measurement area directly on the camera view, receive initial length and width suggestions from the VLM, and confirm or correct them using site information or physical measurements. An overlay preview then displays the area, vehicle trajectories, and estimated speeds together so that the configuration can be checked immediately.</p>
<p>Once calibration is complete, users will be able to choose speed limits and target scenarios for each zone, then connect detected events to EVA alert channels and dashboards. A validated calibration can be reused for a fixed camera, while the interface can request recalibration after the camera position or zoom changes. The central UX goal is to hide unnecessary technical complexity while keeping the basis of the configuration and its results easy to verify.</p>
<p>The vehicle speed estimation technology developed through this research will go through productization—including site-specific setup guidance, VLM-assisted initial values, result-validation views, scenario configuration, and alert integration—and be released as an <strong>official EVA feature.</strong></p>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>Vision Model</category>
        </item>
        <item>
            <title><![CDATA[Observe, Judge, and Question: An Enhanced Thinking Mode]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2</guid>
            <pubDate>Fri, 07 Aug 2026 19:00:00 GMT</pubDate>
            <description><![CDATA[Observe, Judge, and Question: An Enhanced Thinking Mode]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="observe-judge-and-question-an-enhanced-thinking-mode">Observe, Judge, and Question: An Enhanced Thinking Mode<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#observe-judge-and-question-an-enhanced-thinking-mode" class="hash-link" aria-label="Direct link to Observe, Judge, and Question: An Enhanced Thinking Mode" title="Direct link to Observe, Judge, and Question: An Enhanced Thinking Mode">​</a></h2>
<p>It is difficult to expect consistent decisions from a small VLM simply by asking it to “think deeply on its own.” A model may try to solve several problems within a single call, or revisit ambiguous evidence for too long without reaching a conclusion.</p>
<p>In this enhanced Thinking mode, we designed <strong>the reasoning flow itself as a set of explicit roles and structures</strong> instead of leaving every decision to the model’s internal reasoning. The pipeline establishes the criteria first, fixes the area to inspect, evaluates the scene from different perspectives, and leaves the final decision to code.</p>
<p>The goal is not to make the model think more. It is to make the order in which it should <strong>observe, judge, and question</strong> explicit.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-problems-with-the-previous-thinking-structure">1. Problems with the Previous Thinking Structure<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#1-problems-with-the-previous-thinking-structure" class="hash-link" aria-label="Direct link to 1. Problems with the Previous Thinking Structure" title="Direct link to 1. Problems with the Previous Thinking Structure">​</a></h2>
<p>The previous development version judged several hazardous situations in the <code>incidents</code> block with a single VLM call.</p>
<!-- -->
<p>The structure is simple, but assigning different problems to one call creates two issues.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="1-multiple-problems-are-judged-at-once">1) Multiple problems are judged at once<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#1-multiple-problems-are-judged-at-once" class="hash-link" aria-label="Direct link to 1) Multiple problems are judged at once" title="Direct link to 1) Multiple problems are judged at once">​</a></h3>
<p>One call judges <code>fall</code>, <code>smoke</code>, and <code>fire</code> simultaneously. Even when some items are easy to determine, one ambiguous item can consume the entire reasoning budget.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="2-the-scope-of-reasoning-is-difficult-to-control">2) The scope of reasoning is difficult to control<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#2-the-scope-of-reasoning-is-difficult-to-control" class="hash-link" aria-label="Direct link to 2) The scope of reasoning is difficult to control" title="Direct link to 2) The scope of reasoning is difficult to control">​</a></h3>
<p>Reasoning happens inside the model as a black-box process. In practice, the only value the application can adjust is the budget number; it is difficult to control which criteria the model revisits or what it repeatedly checks.</p>
<table><thead><tr><th>Aspect</th><th>Previous structure</th></tr></thead><tbody><tr><td>Targets</td><td>Three fixed types: fall / smoke / fire</td></tr><tr><td>Decision criteria</td><td>Hard-coded in the code</td></tr><tr><td>Evidence</td><td>Not retained; only the final booleans are returned</td></tr><tr><td>Failure behavior</td><td>All results may quietly become False</td></tr></tbody></table>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-limitations-revealed-by-actual-thinking-traces">2. Limitations Revealed by Actual Thinking Traces<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#2-limitations-revealed-by-actual-thinking-traces" class="hash-link" aria-label="Direct link to 2. Limitations Revealed by Actual Thinking Traces" title="Direct link to 2. Limitations Revealed by Actual Thinking Traces">​</a></h2>
<p>Consider a situation in which <code>smoke</code> must be judged in a facility room filled with vapor. A model “thinking deeply” does not necessarily mean that it is finding new evidence.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/trace_smoke_example-97b187c5afc0a7a7d8c481e9be153a3d.png" width="60%"></div>
<p align="center"><i>Facility scene used for the smoke decision</i></p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/trace_qwen_langfuse-d4f120ea20ac6ee3517664313e446b3b.png" width="80%"></div>
<p align="center"><i>Qwen3.5VL-9B-Thinking Langfuse Trace</i></p>
<br>
<p>After finding the white vapor, the model repeatedly reread the same rules to determine whether it was steam or a gas leak. However, no new visual evidence was added during the repetition.</p>
<!-- -->
<p>The repeated thought, “Wait, let’s look closer,” produced <strong>zero new visual evidence</strong>.</p>
<p>The main problems identified in this trace are:</p>
<table><thead><tr><th>Symptom</th><th>Meaning</th></tr></thead><tbody><tr><td>The same rule is reread multiple times</td><td><strong>Rumination</strong> without new information</td></tr><tr><td>The model cycles without reaching a conclusion</td><td>It does not know <strong>how to stop thinking</strong>; the answer depends on where the budget runs out</td></tr><tr><td>The model revisits alternative interpretations from the dataset</td><td>It <strong>reinterprets the policy on the fly</strong> even though the policy was defined in the prompt</td></tr><tr><td>Easy items finish quickly while ambiguous items consume the entire process</td><td>One problem <strong>monopolizes the reasoning budget</strong></td></tr></tbody></table>
<p>As a result, simply asking a small VLM to “think deeply on its own” is not enough. The system needs to guide what the model checks first, when it ends its judgment, and how it handles uncertainty.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-design-philosophy-turn-the-reasoning-flow-into-structure">3. Design Philosophy: Turn the Reasoning Flow into Structure<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#3-design-philosophy-turn-the-reasoning-flow-into-structure" class="hash-link" aria-label="Direct link to 3. Design Philosophy: Turn the Reasoning Flow into Structure" title="Direct link to 3. Design Philosophy: Turn the Reasoning Flow into Structure">​</a></h2>
<p>The enhanced Thinking mode is built around three principles.</p>
<table><thead><tr><th>Principle</th><th>Description</th></tr></thead><tbody><tr><td>1. One call = one problem</td><td>Split the task into small, clear units so that each call does not need separate reasoning. Every call uses an Instruct model with <code>enable_thinking: false</code>.</td></tr><tr><td>2. Reasoning flow = structure</td><td>Guide the flow through the order of establishing criteria → fixing the view → judging → looking for counterevidence.</td></tr><tr><td>3. Code makes the decision</td><td>The final alert is determined by a deterministic Gate, not by a single LLM output.</td></tr></tbody></table>
<p>The new architecture first turns the scenario into explicit rules and then lets voters with different roles make independent judgments.</p>
<!-- -->
<ul>
<li>The two voters run <strong>in parallel and independently</strong>; neither can see the other’s output.</li>
<li>Every call uses an Instruct approach with <code>enable_thinking: false</code>.</li>
<li>Each call receives only one decision problem.</li>
<li>The final alert is triggered only when both voters are positive.</li>
</ul>
<p>The LLM does not decide everything at once. Each stage focuses on its own role, while code controls the connections between stages and the final decision.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-core-mechanism-im-looking-here-now">4. Core Mechanism: “I’m Looking Here Now”<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#4-core-mechanism-im-looking-here-now" class="hash-link" aria-label="Direct link to 4. Core Mechanism: “I’m Looking Here Now”" title="Direct link to 4. Core Mechanism: “I’m Looking Here Now”">​</a></h2>
<p>Before making a judgment, the model is required to output evidence of <strong>where it is looking</strong>. The role responsible for this is <code>focus_locator</code>.</p>
<!-- -->
<p><code>focus_locator</code> does not make the decision. It only suggests the object and location to observe. <code>vote_context</code> then uses both the full frame and a crop of the region specified by the Locator, while <code>vote_skeptic</code> independently verifies the scene from the full frame without relying on the Locator’s description.</p>
<table><thead><tr><th>Mechanism</th><th>Effect</th></tr></thead><tbody><tr><td>Evidence declaration</td><td>Outputs what is visible and where it is located in text and coordinates before making a decision</td></tr><tr><td>Fixed focus</td><td><code>vote_context</code> sees the full frame together with a crop containing a 20% margin around the region</td></tr><tr><td>Contamination prevention</td><td><code>vote_skeptic</code> sees only the full frame without the crop, so an incorrect Locator does not contaminate its judgment</td></tr></tbody></table>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="why-does-this-improve-performance">Why does this improve performance?<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#why-does-this-improve-performance" class="hash-link" aria-label="Direct link to Why does this improve performance?" title="Direct link to Why does this improve performance?">​</a></h3>
<ul>
<li>An 8B-class VLM can easily judge a scene from its overall impression. Providing evidence from an explicit region first helps anchor the judgment to that evidence through <strong>grounding</strong>.</li>
<li>A 20% margin crop enlarges the target while preserving surrounding context. This reduces mistakes caused by looking only at part of the target.</li>
<li>Because the observed location is retained as text, an error can be analyzed as either an incorrect focus or an incorrect judgment.</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-solving-failed-traces-with-structure">5. Solving Failed Traces with Structure<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#5-solving-failed-traces-with-structure" class="hash-link" aria-label="Direct link to 5. Solving Failed Traces with Structure" title="Direct link to 5. Solving Failed Traces with Structure">​</a></h2>
<p>The new Pipeline roles and Gate address the problems found in the previous Thinking traces.</p>
<table><thead><tr><th>Failure in the previous trace</th><th>Response in the new structure</th></tr></thead><tbody><tr><td>Rules are renegotiated during the decision</td><td>Confirm the rules with <code>scenario_rule_builder</code> <strong>before</strong> making the decision</td></tr><tr><td>The inner monologue repeatedly asks, “Is it really smoke?”</td><td>Institutionalize counterevidence through an independent <strong>skeptic voter</strong></td></tr><tr><td>True/False is forced when the state is ambiguous</td><td>A voter can return <code>uncertain</code>, and the Gate blocks the alert</td></tr><tr><td>One problem monopolizes the budget</td><td>Split calls according to the <strong>one call = one problem</strong> principle</td></tr></tbody></table>
<p>This changes the review process required for a decision into observable system stages instead of merely making the model’s internal reasoning longer.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="6-before--after">6. Before / After<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#6-before--after" class="hash-link" aria-label="Direct link to 6. Before / After" title="Direct link to 6. Before / After">​</a></h2>
<table><thead><tr><th>Aspect</th><th>Previous (develop)</th><th>New (thinking_instruct)</th></tr></thead><tbody><tr><td>Targets</td><td>Three fixed types</td><td>Arbitrary scenarios: intrusion, loitering, leakage, and more</td></tr><tr><td>Judgment</td><td>One VLM call judges all items simultaneously</td><td>Four role-separated calls, one problem per call</td></tr><tr><td>Reasoning</td><td>Happens inside the model and is difficult to control</td><td>An Instruct-based structure defines the flow</td></tr><tr><td>Final decision</td><td>Uses the model output directly</td><td>Deterministic AND Gate</td></tr><tr><td>Evidence</td><td>Not separately recorded</td><td>Each voter’s vote and reasoning are recorded</td></tr><tr><td>Decision criteria</td><td>Hard-coded in the code</td><td>Generated dynamically for each scenario</td></tr></tbody></table>
<p>The key change is not simply increasing the number of calls. It separates decision criteria, observation area, positive and counterevidence roles, and the final decision so that failure points can be traced.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="7-more-flexible-scenario-expansion">7. More Flexible Scenario Expansion<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#7-more-flexible-scenario-expansion" class="hash-link" aria-label="Direct link to 7. More Flexible Scenario Expansion" title="Direct link to 7. More Flexible Scenario Expansion">​</a></h2>
<p>In the previous structure, the targets had to be fixed to smoke, fire, and falls because the LLM could easily move outside the intended decision scope.</p>
<p>In the new structure, even when a user edits the <code>incidents</code> block, a <code>rule guide</code> first organizes the decision criteria. Each voter judges only according to the generated criteria.</p>
<p>For example, users can enter requirements such as:</p>
<ul>
<li>“Smoke detected” → <strong>Detect black smoke</strong></li>
<li>“Fire detected” → <strong>Detect large flames occupying most of the region</strong></li>
</ul>
<p>The edited text is used to generate the Rule Guide, and each voter then uses those criteria. This reduces the need to modify the Router or decision code whenever a new scenario is added.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="8-performance-results">8. Performance Results<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#8-performance-results" class="hash-link" aria-label="Direct link to 8. Performance Results" title="Direct link to 8. Performance Results">​</a></h2>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="precision-08273-suppressing-false-positives">Precision 0.8273: Suppressing False Positives<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#precision-08273-suppressing-false-positives" class="hash-link" aria-label="Direct link to Precision 0.8273: Suppressing False Positives" title="Direct link to Precision 0.8273: Suppressing False Positives">​</a></h3>
<p>In the SINGLE evaluation, the new structure substantially improved Precision.</p>
<table><thead><tr><th>Metric</th><th style="text-align:right">F1</th><th style="text-align:right">Accuracy</th><th style="text-align:right"><strong>Precision</strong></th><th style="text-align:right">Recall</th></tr></thead><tbody><tr><td><strong>Before</strong></td><td style="text-align:right">0.8073</td><td style="text-align:right">0.9536</td><td style="text-align:right"><strong>0.7316</strong></td><td style="text-align:right">0.9006</td></tr><tr><td><strong>After</strong></td><td style="text-align:right">0.8042</td><td style="text-align:right">0.9518</td><td style="text-align:right"><strong>0.8273</strong></td><td style="text-align:right">0.7823</td></tr></tbody></table>
<p>Precision increased from 0.7316 to 0.8273, a gain of 0.0957 points. Recall decreased from 0.9006 to 0.7823. The AND Gate, which requires both voters to be positive before triggering an alert, suppresses false positives at the cost of conservatively blocking some true cases.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="category-level-performance">Category-level performance<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#category-level-performance" class="hash-link" aria-label="Direct link to Category-level performance" title="Direct link to Category-level performance">​</a></h3>
<table><thead><tr><th>Category</th><th style="text-align:right">Before Precision</th><th style="text-align:right">Before Recall</th><th style="text-align:right">Before F1</th><th style="text-align:right">After Precision</th><th style="text-align:right">After Recall</th><th style="text-align:right">After F1</th></tr></thead><tbody><tr><td>Smoke / flames</td><td style="text-align:right">1.0000</td><td style="text-align:right">0.9714</td><td style="text-align:right">0.9855</td><td style="text-align:right">1.0000</td><td style="text-align:right">1.0000</td><td style="text-align:right">1.0000</td></tr><tr><td>Fire</td><td style="text-align:right">0.7288</td><td style="text-align:right">0.8600</td><td style="text-align:right">0.7890</td><td style="text-align:right"><strong>0.8085</strong></td><td style="text-align:right">0.7600</td><td style="text-align:right">0.7835</td></tr><tr><td>Fall</td><td style="text-align:right">0.5319</td><td style="text-align:right">0.8621</td><td style="text-align:right">0.6579</td><td style="text-align:right"><strong>0.7222</strong></td><td style="text-align:right">0.6610</td><td style="text-align:right">0.6903</td></tr></tbody></table>
<p>Smoke and flames reached 1.0000 across all metrics, while Fire Precision improved to 0.8085. In particular, false fire alerts were reduced to nine cases, reducing unnecessary alerts in operations.</p>
<p>In fall scenarios such as the example above, the system checks step by step whether a person is lying horizontally on the floor rather than simply crouching or sitting, and whether there is sufficient direct visual evidence.</p>
<p>Conversely, a red glow is not immediately treated as a real flame. The system reviews the lighting, reflections, flame shape, combustion clues, and whether smoke is present; without clear evidence, it returns <code>alert: false</code>.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="summary">Summary<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode_v2#summary" class="hash-link" aria-label="Direct link to Summary" title="Direct link to Summary">​</a></h2>
<blockquote>
<strong>Instead of asking an 8B model to think more deeply, we laid out the path that its reasoning should follow — establish criteria, fix the focus, judge, and question.</strong>
</blockquote>
<br>
<hr>
<br>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>Vision Model</category>
            <category>AI Agent</category>
        </item>
        <item>
            <title><![CDATA[From Router-Based Agents to Tool-Calling Agents]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat</guid>
            <pubDate>Fri, 07 Aug 2026 18:00:00 GMT</pubDate>
            <description><![CDATA[From Router-Based Agents to Tool-Calling Agents]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="from-router-based-agents-to-tool-calling-agents">From Router-Based Agents to Tool-Calling Agents<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#from-router-based-agents-to-tool-calling-agents" class="hash-link" aria-label="Direct link to From Router-Based Agents to Tool-Calling Agents" title="Direct link to From Router-Based Agents to Tool-Calling Agents">​</a></h2>
<p>As EVA's features and user requests grow, deciding which predefined route should handle each request becomes increasingly difficult. In EVA v3.1, we changed the chat architecture from a router-based agent to a tool-calling agent.</p>
<p>The previous architecture classified a request and sent it to a predefined processing route. The new architecture gives the LLM a registry of available capabilities, allowing it to select the tool required by the request and respond based on its execution result.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-why-change-the-architecture">1. Why change the architecture?<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#1-why-change-the-architecture" class="hash-link" aria-label="Direct link to 1. Why change the architecture?" title="Direct link to 1. Why change the architecture?">​</a></h2>
<p>A router-based architecture has a clear advantage: each request type has an explicit processing path. However, as the number of features and question types increases, several limitations become visible:</p>
<ul>
<li>Adding a feature often requires changes to both the router and the graph branches connected to it.</li>
<li>Manual content embedded in prompts and code becomes difficult to update and maintain.</li>
<li>Questions about the current runtime configuration can become mixed with general feature explanations in the same Q&amp;A flow.</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-router-based-agents-and-tool-calling-agents">2. Router-based agents and tool-calling agents<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#2-router-based-agents-and-tool-calling-agents" class="hash-link" aria-label="Direct link to 2. Router-based agents and tool-calling agents" title="Direct link to 2. Router-based agents and tool-calling agents">​</a></h2>
<p>The two architectures answer different questions.</p>
<table><thead><tr><th>Architecture</th><th>Main question</th><th>Typical flow</th></tr></thead><tbody><tr><td>Router-based agent</td><td>“Which path should handle this request?”</td><td>Request → classification → predefined Node or Subgraph → response</td></tr><tr><td>Tool-calling agent</td><td>“Which capability is needed to handle this request?”</td><td>Request → intent analysis → required tool selection → tool execution → response</td></tr></tbody></table>
<p>In the router-based design, the code defines the available branches and selects one of them. In the tool-calling design, the LLM receives the available tool schemas and determines the tool required by the request.</p>
<p>The distinction is not simply about replacing a router with an LLM. It changes where the processing decision is made and how capabilities are added to the system:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Router-based agent</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">User request</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Request-type classification</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → One predefined route</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Node or Subgraph execution</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Response</span><br></span></code></pre></div></div>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Tool-calling agent</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">User request</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Intent analysis</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Required tool selected</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Tool execution</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Execution result reviewed</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Response</span><br></span></code></pre></div></div>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-the-preparedecidetoolcomposefinalize-flow">3. The Prepare–Decide–Tool–Compose–Finalize flow<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#3-the-preparedecidetoolcomposefinalize-flow" class="hash-link" aria-label="Direct link to 3. The Prepare–Decide–Tool–Compose–Finalize flow" title="Direct link to 3. The Prepare–Decide–Tool–Compose–Finalize flow">​</a></h2>
<p>The new agent follows five conceptual stages:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">User request</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Prepare</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Decide</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Tool</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Compose</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Finalize</span><br></span></code></pre></div></div>
<table><thead><tr><th>Stage</th><th>Role</th></tr></thead><tbody><tr><td><code>Prepare</code></td><td>Prepare the language, conversation history, current settings, and other context required for the request.</td></tr><tr><td><code>Decide</code></td><td>Analyze the request and select the required tool.</td></tr><tr><td><code>Tool</code></td><td>Execute the selected tool.</td></tr><tr><td><code>Compose</code></td><td>Turn the tool execution result into a natural response.</td></tr><tr><td><code>Finalize</code></td><td>Convert the result into the final response format.</td></tr></tbody></table>
<p>For example, the request “Change the AI inference interval to 30 seconds” can be processed as follows:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Decide</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Select set_detection_interval</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Validate the input value: 30 seconds</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Change the setting</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Return the execution result</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Generate the final response</span><br></span></code></pre></div></div>
<p>Not every request needs every stage to perform a separate operation. The important point is that the agent has a consistent lifecycle in which tool selection, execution, and response generation are explicit.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-moving-manual-qa-from-prompts-to-knowledge-rag">4. Moving manual Q&amp;A from prompts to Knowledge RAG<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#4-moving-manual-qa-from-prompts-to-knowledge-rag" class="hash-link" aria-label="Direct link to 4. Moving manual Q&amp;A from prompts to Knowledge RAG" title="Direct link to 4. Moving manual Q&amp;A from prompts to Knowledge RAG">​</a></h2>
<p>The architectural change also affects how EVA answers questions about its manuals and features.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="prompt-based-chat">Prompt-based Chat<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#prompt-based-chat" class="hash-link" aria-label="Direct link to Prompt-based Chat" title="Direct link to Prompt-based Chat">​</a></h3>
<p>In the previous design, guides were selected by request type and inserted into the prompt before generating an answer.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">User question</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Classify the question</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Select a related guide</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Include the guide in the prompt</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Generate the LLM response</span><br></span></code></pre></div></div>
<p>For example, a question about object-detection sensitivity might be classified as a terminology or app-feature question. The code would then select <code>TERM_GUIDE</code> or <code>APP_GUIDE</code> and include its content in the prompt.</p>
<p>This approach has several limitations:</p>
<ul>
<li>There is no separate document-retrieval step.</li>
<li>The answer is limited to the guide content included in the prompt.</li>
<li>Updating a guide may require changes to a prompt or code.</li>
<li>The code must select the guide that matches the question type.</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="tool-calling-chat">Tool-calling Chat<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#tool-calling-chat" class="hash-link" aria-label="Direct link to Tool-calling Chat" title="Direct link to Tool-calling Chat">​</a></h3>
<p>In the new design, a manual question selects <code>answer_eva_question</code>. The tool embeds the question, searches the Knowledge store, and provides relevant pages as context for the answer model.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">User question</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Select answer_eva_question</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Create a question embedding</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Search related pages in Qdrant</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Pass the retrieved page content as context</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Generate the answer</span><br></span></code></pre></div></div>
<p>For a question such as “How do I configure a detection scenario in EVA?”, the system searches for the manual page containing the relevant configuration procedure and generates an answer from that page.</p>
<p>This makes manuals an independently managed Knowledge source. Only the relevant pages can be retrieved, and metadata such as the file name and page number can be included in the result.</p>
<p>The trade-off is that retrieval quality now directly affects answer quality. The Knowledge pipeline and its evaluation become important parts of the chat system.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-page-level-knowledge-ingest">5. Page-level Knowledge Ingest<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#5-page-level-knowledge-ingest" class="hash-link" aria-label="Direct link to 5. Page-level Knowledge Ingest" title="Direct link to 5. Page-level Knowledge Ingest">​</a></h2>
<p>RAG search requires manuals to be stored in a searchable form before users ask questions. The current ingestion flow is:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">PDF or Markdown</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Extract text, tables, and figure information</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Build page-level content</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Create embeddings</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Store page vectors in Qdrant</span><br></span></code></pre></div></div>
<p>At query time, the flow is:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">User question</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Create a query embedding</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Search Qdrant by similarity</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Return related pages</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Use the pages as answer context</span><br></span></code></pre></div></div>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="why-search-by-page">Why search by page?<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#why-search-by-page" class="hash-link" aria-label="Direct link to Why search by page?" title="Direct link to Why search by page?">​</a></h3>
<p>A page often contains information that is meaningful only when read together. For example, a single page may include:</p>
<ul>
<li>A feature description at the top.</li>
<li>A configuration screen image in the middle.</li>
<li>Important cautions at the bottom.</li>
</ul>
<p>If the page is split into very small sentence-level chunks, these related pieces can be returned as separate search results. Page-level retrieval preserves the relationship between the description, configuration procedure, table, image explanation, and caution.</p>
<p>The goal is not simply to make chunks larger. It is to preserve information that is explained together within the same page.</p>
<p>The current structure finds related material at page granularity and passes the page text and metadata to the answer model as a ToolResult. This preserves the relationship between descriptions, configuration procedures, tables, and cautions while separating Knowledge retrieval from answer generation.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="extending-knowledge-for-customer-environments">Extending Knowledge for customer environments<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#extending-knowledge-for-customer-environments" class="hash-link" aria-label="Direct link to Extending Knowledge for customer environments" title="Direct link to Extending Knowledge for customer environments">​</a></h3>
<p>Knowledge can include not only EVA manuals but also documents required for each customer environment. Customer operating manuals, equipment standards, work procedures, and safety and environmental regulations can be ingested so that answers are grounded in the standards and rules of the relevant site, rather than limited to general feature explanations.</p>
<p>For example, adding a safety and environmental regulations handbook to Knowledge makes it possible to ask not only whether a detection alarm occurred, but also which safety or environmental regulations apply, what the site response procedure is, and what follow-up reporting is required. Customer-specific documents can be managed as separate Knowledge sources while using the same tool-calling flow to search and answer according to each site’s operating rules.</p>
<br>
<hr>
<br>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="6-end-to-end-architecture">6. End-to-end architecture<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#6-end-to-end-architecture" class="hash-link" aria-label="Direct link to 6. End-to-end architecture" title="Direct link to 6. End-to-end architecture">​</a></h2>
<p>The complete system has two related flows: an offline Knowledge ingestion flow and an online inference flow.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Offline Knowledge flow</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">PDF / Markdown</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Extract text, tables, and figures</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Create page-level text</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Generate page embeddings</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Store vectors in Qdrant</span><br></span></code></pre></div></div>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Online inference flow</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">User request</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Prepare conversation history and current settings</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Bind the available Tool Registry to the LLM</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → LLM creates the required tool call</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Execute the selected tool</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Return the ToolResult</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Generate a natural-language answer when needed</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Finalize the response</span><br></span></code></pre></div></div>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="tool-calling-chat-1">Tool-calling Chat<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#tool-calling-chat-1" class="hash-link" aria-label="Direct link to Tool-calling Chat" title="Direct link to Tool-calling Chat">​</a></h3>
<!-- -->
<p>The Tool Registry can contain different kinds of capabilities:</p>
<ul>
<li>Action Tools for changing application state.</li>
<li>Q&amp;A Tools for Knowledge-backed answers.</li>
<li>Conversation Tools for general dialogue.</li>
<li>Meta Tools for reading runtime information or other system state.</li>
</ul>
<p>When <code>answer_eva_question</code> is selected, it starts the query-embedding and Qdrant-search flow. The retrieved page content is returned as a ToolResult and can then be used by the composing model.</p>
<p>This keeps the responsibilities separate: tools perform concrete operations, Knowledge retrieval provides evidence, and the language model generates the final response.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="7-what-the-structural-change-enables">7. What the structural change enables<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#7-what-the-structural-change-enables" class="hash-link" aria-label="Direct link to 7. What the structural change enables" title="Direct link to 7. What the structural change enables">​</a></h2>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="the-practical-benefits-of-tool-calling">The practical benefits of tool calling<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#the-practical-benefits-of-tool-calling" class="hash-link" aria-label="Direct link to The practical benefits of tool calling" title="Direct link to The practical benefits of tool calling">​</a></h3>
<p>The value of tool calling is not simply that an LLM can call a function. Its main benefit is that request interpretation, concrete operations, Knowledge retrieval, and final answer generation have separate responsibilities and clear roles.</p>
<ul>
<li><strong>Easier feature expansion</strong>: Add a tool and its schema to the registry instead of wiring every new capability into router branches and graph paths.</li>
<li><strong>Current information</strong>: Retrieve manuals and runtime state from Knowledge or state-lookup tools when needed instead of fixing them inside a prompt.</li>
<li><strong>Safer execution</strong>: Keep natural-language interpretation separate from state changes by applying validation, permissions, and execution conditions to each Action Tool.</li>
<li><strong>Better observability</strong>: Record and evaluate which tool was selected, which inputs were passed, and which search results were returned.</li>
</ul>
<p>For the evaluated requests, the average response speed was about <strong>30% faster</strong> than the previous router-based flow. The improvement came from reducing unnecessary classification steps and executing the required tool directly.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="easier-feature-expansion">Easier feature expansion<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#easier-feature-expansion" class="hash-link" aria-label="Direct link to Easier feature expansion" title="Direct link to Easier feature expansion">​</a></h3>
<p>In the previous architecture, a new capability could require changes to router and graph branches. In the new architecture, the capability can be implemented as a tool and registered with its schema.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Before</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Modify the router and graph branches</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">After</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Implement a tool</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Register its schema</span><br></span></code></pre></div></div>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="independent-knowledge-management">Independent Knowledge management<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#independent-knowledge-management" class="hash-link" aria-label="Direct link to Independent Knowledge management" title="Direct link to Independent Knowledge management">​</a></h3>
<p>Manual content no longer needs to live in prompts or application code. It can be managed as Knowledge documents and reflected through the ingestion pipeline.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="clearer-separation-of-request-purposes">Clearer separation of request purposes<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#clearer-separation-of-request-purposes" class="hash-link" aria-label="Direct link to Clearer separation of request purposes" title="Direct link to Clearer separation of request purposes">​</a></h3>
<p>The same user-facing topic can be connected to different backends depending on the request:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Current runtime value</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Runtime setting lookup</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Feature explanation</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Knowledge search</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Configuration change</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  → Action Tool</span><br></span></code></pre></div></div>
<p>Along with the response-speed improvement, the processing path becomes more flexible as the number of capabilities grows.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="8-new-operational-considerations">8. New operational considerations<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#8-new-operational-considerations" class="hash-link" aria-label="Direct link to 8. New operational considerations" title="Direct link to 8. New operational considerations">​</a></h2>
<p>The architecture is more flexible, but it also introduces new areas that require monitoring and evaluation.</p>
<table><thead><tr><th>Area</th><th>Question to answer</th></tr></thead><tbody><tr><td>Tool selection</td><td>Does the agent select the correct tool for the request?</td></tr><tr><td>Tool schema</td><td>Can the LLM correctly understand the tool's purpose and input values?</td></tr><tr><td>Action safety</td><td>Which state-changing actions can be executed automatically?</td></tr><tr><td>RAG retrieval</td><td>Does the search return the relevant page without missing key information?</td></tr><tr><td>Page-level retrieval</td><td>Does one page contain too many unrelated topics?</td></tr><tr><td>Document management</td><td>How are versions and duplicate Qdrant points managed?</td></tr><tr><td>Ingestion</td><td>Do the service settings and ingestion script remain consistent?</td></tr><tr><td>Evaluation</td><td>How are tool selection, action execution, and RAG answers tested?</td></tr></tbody></table>
<p>In particular, a successful tool-calling system needs to evaluate more than the final wording of the answer. It must also inspect whether the right tool was selected, whether its inputs were valid, whether the action was safe, and whether the answer was grounded in the right Knowledge page.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="closing">Closing<a href="https://spectrabrain.ai/en/blog_tech/Research/tool_calling_chat#closing" class="hash-link" aria-label="Direct link to Closing" title="Direct link to Closing">​</a></h2>
<p>The move from a router-based agent to a tool-calling agent is not simply a replacement of one chat component. It is a change from a system where code determines the request's processing path to a system where the LLM selects the capabilities and Knowledge required to handle the request.</p>
<p>With this structure, EVA can manage manuals as an independent Knowledge source and extend functionality by adding tools to a registry. At the same time, tool schemas, action safety, retrieval quality, document versioning, and end-to-end evaluation become essential parts of operating the agent reliably.</p>
<br>
<hr>
<br>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>AI Agent</category>
            <category>rag</category>
        </item>
        <item>
            <title><![CDATA[EVA Scope: Observability for Large-Scale CCTV AI Operations]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Innovation/evascope</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Innovation/evascope</guid>
            <pubDate>Thu, 06 Aug 2026 18:00:00 GMT</pubDate>
            <description><![CDATA[EVA Scope uses Frame Journey data generated by EVA to observe the entire AI pipeline and analyze bottlenecks, latency, idle time, and processing imbalance.]]></description>
            <content:encoded><![CDATA[<p>In large-scale CCTV AI operations, the stability and efficiency of EVA cannot be assessed with a single metric. When dozens of EVAs and hundreds or thousands of cameras are in operation, each frame passes through multiple processing stages inside the App after entering from a camera. This may be manageable when the number of cameras or EVAs is small, but as the operation grows, it becomes increasingly difficult to verify that every frame is following its intended path.</p>
<p>What matters is not simply whether every frame was processed. Operators also need to understand which processing stages are consuming resources, where resources are sitting idle, and whether processing is concentrated on particular cameras or Connections.</p>
<p>EVA Scope observes the end-to-end processing flow of the AI pipeline using Frame Journey data generated by EVA. It provides a comprehensive view of bottlenecks, latency, idle time, and processing imbalance, giving operators the evidence they need to determine how EVA can operate more reliably and efficiently within a limited operating environment.</p>
<p>This article introduces the architecture and key capabilities of EVA Scope, designed to observe frame-processing status in large-scale CCTV AI operations.</p>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-eva-scope-at-a-glance">1. EVA Scope at a Glance<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#1-eva-scope-at-a-glance" class="hash-link" aria-label="Direct link to 1. EVA Scope at a Glance" title="Direct link to 1. EVA Scope at a Glance">​</a></h2>
<p><img decoding="async" loading="lazy" alt="Connections List overview" src="https://spectrabrain.ai/en/assets/images/Connections-b814474550a880321ad3fe995f3e2c36.png" width="1710" height="1307" class="img_ev3q"></p>
<p><em>View the status and Risk Score of numerous Connections on a single screen and prioritize the ones that need attention.</em></p>
<p>Connections List provides the following information for each Connection:</p>
<ul>
<li>Normal, caution, and warning status</li>
<li>Latest Risk Score</li>
<li>Bottleneck, Idle, Imbalance, and Error indexes</li>
<li>Recent hourly status trends</li>
<li>Total frame throughput</li>
<li>Latest data collection status</li>
<li>EVA App connection status</li>
<li>Number of active Alerts</li>
</ul>
<p>Operators do not need to inspect every Connection in sequence. They can start with the Connections that have the highest risk, check when the status began to change, and then move to the detailed analysis screen.</p>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-signal-to-insight-flow">2. Signal-to-Insight Flow<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#2-signal-to-insight-flow" class="hash-link" aria-label="Direct link to 2. Signal-to-Insight Flow" title="Direct link to 2. Signal-to-Insight Flow">​</a></h2>
<p>The core of EVA Scope is not simply displaying a large amount of data. It identifies signals that affect processing efficiency in a large volume of frame data and connects them to analysis at the Connection, camera, processing-stage, and individual-frame levels.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="connections-list">Connections List<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#connections-list" class="hash-link" aria-label="Direct link to Connections List" title="Direct link to Connections List">​</a></h3>
<p>The first step is to review the list and status of all Connections, or EVA Apps. EVA Scope explains the factors that affect processing efficiency through four detailed indexes and sets Connection-level priorities with the Risk Score.</p>
<p>Operators can answer questions such as:</p>
<ul>
<li>Which Connection has the highest Risk Score?</li>
<li>Is recent processing efficiency improving or deteriorating?</li>
<li>Did throughput or latency change during a particular time period?</li>
<li>Is data collection stable for a particular Connection?</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="alert-history">Alert History<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#alert-history" class="hash-link" aria-label="Direct link to Alert History" title="Direct link to Alert History">​</a></h3>
<p>When a key metric moves outside its configured threshold or the data collection status requires additional attention, the event is recorded in Alert History. Alert History displays the time period, Connection, status, and Risk Score associated with the problem.</p>
<p><img decoding="async" loading="lazy" alt="Alert History" src="https://spectrabrain.ai/en/assets/images/Alerthistory-9d89c5df104656214ed66d4b832aee87.png" width="2096" height="1300" class="img_ev3q"></p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="hourly-overview">Hourly Overview<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#hourly-overview" class="hash-link" aria-label="Direct link to Hourly Overview" title="Direct link to Hourly Overview">​</a></h3>
<p>Select the Connection where the problem occurred. By setting a start and end time, operators can analyze the Image Frames that entered during that period. They can review the frames’ end states and status trends, processing times, and changes in key metrics to identify the problem. Historical periods can also be analyzed within the seven-day retention window.</p>
<p><img decoding="async" loading="lazy" alt="Hourly Overview" src="https://spectrabrain.ai/en/assets/images/HourlyOverview-34eade5affc9c2a2dd83c15f8bdfc762.png" width="2096" height="1300" class="img_ev3q"></p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="index-analysis">Index Analysis<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#index-analysis" class="hash-link" aria-label="Direct link to Index Analysis" title="Direct link to Index Analysis">​</a></h3>
<p>EVA Scope provides four indexes that make up processing efficiency. Because it shows changes in each index together with its detailed components, operators can identify the factors behind a high metric rather than simply observing that the metric is high.</p>
<ul>
<li>Bottleneck Index: Processing delays and Queue waiting</li>
<li>Idle Index: Idle levels in the Vision, Agent, and frame-ingestion stages</li>
<li>Imbalance Index: Processing imbalance across cameras</li>
<li>Error Index: Ratio of erroneous frames</li>
</ul>
<p><img decoding="async" loading="lazy" alt="Index Analysis" src="https://spectrabrain.ai/en/assets/images/IndexAnalysis-c79fb2dd8c67cd25462207ac18b3efa0.png" width="1710" height="1307" class="img_ev3q"></p>
<p>This flow allows operators to go beyond checking an Alert and determine which stages and resources should be adjusted first to improve processing efficiency. EVA Scope does not decide the improvement plan on its own. Instead, it provides the evidence needed to consider follow-up actions such as reallocating resources, adjusting Queue policies, reviewing camera-specific processing policies, and improving the inference environment.</p>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-why-observability-is-needed-in-large-scale-cctv-operations">3. Why Observability Is Needed in Large-Scale CCTV Operations<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#3-why-observability-is-needed-in-large-scale-cctv-operations" class="hash-link" aria-label="Direct link to 3. Why Observability Is Needed in Large-Scale CCTV Operations" title="Direct link to 3. Why Observability Is Needed in Large-Scale CCTV Operations">​</a></h2>
<p>In a small-scale environment, checking a few cameras and a single EVA App can provide a reasonable understanding of the system status.</p>
<p>As the operation grows, however, the following problems emerge:</p>
<ul>
<li>The status of dozens of EVA Apps must be checked individually.</li>
<li>Frame-processing status is distributed across hundreds or thousands of cameras.</li>
<li>The final Drop result does not reveal where a failure occurred.</li>
<li>It is difficult to connect the time of a problem to the original frame.</li>
<li>Users must find the relevant data manually after noticing a problem.</li>
</ul>
<p>An image-frame processing pipeline is not a simple request-response structure.</p>
<p>A single frame passes through multiple asynchronous stages. Some frames may be Dropped during processing or end in a different state. As a result, alerts visible in the UI alone are not enough to determine whether EVA is operating in an optimal state.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Frame Ingest</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Frame Queue</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Vision Inference</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Agent Queue</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Agent Inference</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Finished / Dropped / Error</span><br></span></code></pre></div></div>
<p>To determine whether EVA is operating optimally, operators need to know which path each frame took and how much time it spent at each stage.</p>
<p>EVA Scope combines three types of information:</p>
<ol>
<li>The final state of the frame</li>
<li>Timestamps for each processing stage</li>
<li>Processing context by Connection, camera, and time period</li>
</ol>
<p>This makes it possible to answer questions beyond simply saying that inefficiency occurred:</p>
<ul>
<li>Which EVA App experienced the bottleneck?</li>
<li>At what time did the load begin to increase?</li>
<li>Is the bottleneck in Vision, Agent, or the Queue?</li>
<li>Is the problem concentrated on a particular camera?</li>
<li>Is the bottleneck still occurring?</li>
<li>Did processing efficiency actually change after an action was taken?</li>
</ul>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-eva-scope-core-intelligence">4. EVA Scope Core Intelligence<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#4-eva-scope-core-intelligence" class="hash-link" aria-label="Direct link to 4. EVA Scope Core Intelligence" title="Direct link to 4. EVA Scope Core Intelligence">​</a></h2>
<p>EVA Scope Core Intelligence is the core technology layer that collects and aggregates large volumes of frame data and transforms them into signals that can be used for operations.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="41-overall-architecture">4.1 Overall Architecture<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#41-overall-architecture" class="hash-link" aria-label="Direct link to 4.1 Overall Architecture" title="Direct link to 4.1 Overall Architecture">​</a></h3>
<p>EVA Scope collects Image Frame Journey data from EVA Apps, stores it in a database, and analyzes it through a dashboard. In EVA Scope, a Journey refers to the processing path of an image frame from Ingest through Queue, Vision, and Agent to its final state.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Dozens of EVA Apps</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Hundreds to thousands of Cameras</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">EVA App stats-dump (JSONL)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">FastAPI Backend</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Streaming Parser</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Hourly Rollup Accumulator</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">ClickHouse</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ├─ Raw Frame Data</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ├─ Hourly Health Rollup</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ├─ Ingestion Ledger</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    └─ Alerts</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Next.js Dashboard</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ├─ Connections List</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ├─ Alert History</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    └─ Detailed Analysis</span><br></span></code></pre></div></div>
<p>The current design considers dump files of approximately 50 MB per Connection per hour and a maximum of 100 Connections. At this scale, it is important to separate operational monitoring data from detailed analysis data instead of querying all Raw data every time.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="42-raw-and-rollup-data-collection-and-storage-strategy-at-scale">4.2 Raw and Rollup: Data Collection and Storage Strategy at Scale<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#42-raw-and-rollup-data-collection-and-storage-strategy-at-scale" class="hash-link" aria-label="Direct link to 4.2 Raw and Rollup: Data Collection and Storage Strategy at Scale" title="Direct link to 4.2 Raw and Rollup: Data Collection and Storage Strategy at Scale">​</a></h3>
<p>When all Raw frame data is stored and queried continuously in a large-scale environment, storage requirements and query costs increase rapidly. EVA Scope processes data as a stream and separates Raw data from Rollup data during collection. With the <code>rollup_only</code> policy, the stream is consumed to create hourly aggregations, while Raw frames are not stored. With the <code>full</code> policy, the same stream is used for both Rollup aggregation and Raw batch inserts.</p>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="streaming-based-data-processing">Streaming-Based Data Processing<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#streaming-based-data-processing" class="hash-link" aria-label="Direct link to Streaming-Based Data Processing" title="Direct link to Streaming-Based Data Processing">​</a></h4>
<p>EVA App statistics are provided in JSONL format. EVA Scope reads the dump line by line and aggregates only the required information instead of storing the entire dump in the database or loading it into memory.</p>
<ul>
<li>Frame counts by status</li>
<li>Frame counts and Drop counts by camera</li>
<li>End-to-End latency</li>
<li>Vision and Agent inference time</li>
<li>Queue waiting ratio</li>
<li>Actual time spent at each processing stage</li>
</ul>
<p>This approach prevents memory usage from growing in proportion to the total size of the dump file.</p>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="rollup-data">Rollup Data<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#rollup-data" class="hash-link" aria-label="Direct link to Rollup Data" title="Direct link to Rollup Data">​</a></h4>
<p>Rollup is operational data aggregated by Connection and time bucket. It provides only the information needed for operations, processed from the data collected through streaming.</p>
<ul>
<li>Total frame count</li>
<li>Finished, Filtered, Dropped, and Error counts</li>
<li>End-to-End P50/P95 latency</li>
<li>Frame and Drop counts by camera</li>
<li>Bottleneck, Idle, and Imbalance indexes</li>
<li>Risk Score</li>
</ul>
<p>Connections List and Alert detection use Rollup data. This allows operators to quickly check the overall operational status without scanning thousands of cameras at the Raw-data level every time.</p>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="raw-data">Raw Data<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#raw-data" class="hash-link" aria-label="Direct link to Raw Data" title="Direct link to Raw Data">​</a></h4>
<p>Raw data is used for detailed analysis of individual frames. EVA Scope does not collect all Raw data; it collects and stores Raw data only for the Connection and time period that an operator wants to analyze.</p>
<ul>
<li>Inspecting error frames from a specific camera</li>
<li>Tracing the processing path of an individual frame</li>
<li>Checking Vision and Agent processing time</li>
<li>Detailed analysis of a specific time period</li>
</ul>
<p>EVA Scope also provides Connection-level storage policies.</p>
<table><thead><tr><th>Policy</th><th>Description</th></tr></thead><tbody><tr><td><code>rollup_only</code></td><td>Default policy. Stores aggregated data only</td></tr><tr><td><code>full</code></td><td>Stores Raw data continuously</td></tr></tbody></table>
<p>Under the default policy, operations are centered on Rollup data. When an operator requests detailed analysis, Raw data for the required time period can be collected again. In other words, EVA Scope does not store all data in the same way. Operational data is aggregated lightly for fast queries, while data needed for root-cause analysis is separated so that it can be inspected in detail when needed. With this Rollup policy, EVA Scope maintains monitoring capabilities for all Connections while reducing storage requirements by 99%.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="43-key-metrics-design">4.3 Key Metrics Design<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#43-key-metrics-design" class="hash-link" aria-label="Direct link to 4.3 Key Metrics Design" title="Direct link to 4.3 Key Metrics Design">​</a></h3>
<p>EVA Scope uses a range of metrics to diagnose operational efficiency and effectiveness in large-scale environments.</p>
<table><thead><tr><th>Metric</th><th>Meaning</th></tr></thead><tbody><tr><td>Bottleneck Index</td><td>Risk of frame loss caused by processing delays and Queue waiting</td></tr><tr><td>Idle Index</td><td>Idle level in the Vision, Agent, and camera-ingestion stages</td></tr><tr><td>Imbalance Index</td><td>Degree to which processing or Drops are concentrated on particular cameras</td></tr><tr><td>Error Index</td><td>Ratio of erroneous frames among all frames</td></tr><tr><td>Risk Score</td><td>Connection-level risk calculated from multiple metrics</td></tr></tbody></table>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="closing">Closing<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evascope#closing" class="hash-link" aria-label="Direct link to Closing" title="Direct link to Closing">​</a></h2>
<p>The real challenge in a large-scale CCTV AI environment is not a lack of infrastructure but the opacity of the AI system’s internal behavior. If operators do not know where bottlenecks occur or which resources are sitting idle among the frames arriving from thousands of cameras, adding GPU servers will not solve the problem. Previously, it was difficult even to identify which EVA among the many running EVAs needed improvement. Even after finding one, analyzing system logs could take tens or hundreds of minutes. EVA Scope dramatically shortens this process by converting it into intuitive signals. By visualizing the entire Journey from image-frame ingestion to the final Agent inference, it enables operators to optimize the configuration and settings of the AI system based on clear data.</p>
<blockquote>
<p><em>Beyond monitoring that simply checks whether “the server is alive,”
and beyond an operations system that checks “where a failure occurred,”
EVA Scope enables operators to determine what to observe and what to improve so that EVA can operate stably and efficiently within limited system resources.</em></p>
</blockquote>
<p>This is the value that EVA Scope aims to provide in large-scale CCTV AI operations.</p>]]></content:encoded>
            <category>Tech</category>
            <category>EVA</category>
            <category>AI Agent</category>
            <category>observability</category>
        </item>
        <item>
            <title><![CDATA[EVA mentor: Don’t Just Watch AI in the Field—Teach It]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Innovation/evamentor</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Innovation/evamentor</guid>
            <pubDate>Fri, 31 Jul 2026 18:00:00 GMT</pubDate>
            <description><![CDATA[EVAmentor uses EVA-generated alerts, feedback, and review data to diagnose performance, identify the causes of false positives and directions for improvement, and validate proposed changes with A/B tests before deployment. This lets operators improve the most important issues without checking hundreds of cameras one by one, while continuously operating reliable AI detection tuned to each site.]]></description>
            <content:encoded><![CDATA[<p>EVA detects dangerous situations such as fire and smoke, worker falls, and missing personal protective equipment from CCTV footage and generates alerts. However, even the same scenario can produce very different detection results and false-positive patterns depending on the site and camera environment.</p>
<p>As the number of cameras and sites grows, the volume of alerts grows with it. When false positives accumulate, field operators stop trusting the alerts. Checking only whether cameras are connected or whether the system is running is not enough to improve detection quality; operators also need to know which cameras and scenarios need attention and whether a configuration change actually worked.</p>
<p><strong>EVAmentor</strong> is an operational solution that measures performance from EVA decisions and user feedback, diagnoses problems, finds directions for improvement, and validates the results. It goes beyond monitoring system status by providing mentoring that helps operators understand the causes of false positives and decide which feedback and settings to apply. EVAmentor’s role is to help continuously refine detection quality.</p>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-evamentor-at-a-glance">1. EVAmentor at a Glance<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#1-evamentor-at-a-glance" class="hash-link" aria-label="Direct link to 1. EVAmentor at a Glance" title="Direct link to 1. EVAmentor at a Glance">​</a></h2>
<p>EVAmentor connects the operational cycle of <strong>finding problems → diagnosing causes → improving → validating</strong> in one flow. Operators can review alerts, feedback, and configuration-change history using the same criteria and decide what to check and improve first.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="overview--the-first-screen-for-understanding-detection-status">Overview — The First Screen for Understanding Detection Status<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#overview--the-first-screen-for-understanding-detection-status" class="hash-link" aria-label="Direct link to Overview — The First Screen for Understanding Detection Status" title="Direct link to Overview — The First Screen for Understanding Detection Status">​</a></h3>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/image-9ac3eedf4662129f09848dba29ac0907.png" width="100%"></div>
<p>For each site, operators can see the number of cameras, alert load, alert accuracy, feedback effectiveness, and recent alerts at a glance. The screen shows not only the number of alerts, but also the share judged to be real risks and how much user feedback contributes to reducing false positives. This makes it easier to understand the current detection quality and identify what needs attention first.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="action-required--deciding-what-to-review-first">Action Required — Deciding What to Review First<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#action-required--deciding-what-to-review-first" class="hash-link" aria-label="Direct link to Action Required — Deciding What to Review First" title="Direct link to Action Required — Deciding What to Review First">​</a></h3>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/image-1-feb2ad2affaf70bd32962a3d0328f345.png" width="100%"></div>
<p>Instead of manually checking dozens or hundreds of cameras and scenarios, EVAmentor automatically classifies their status based on recent data and review history. Operators can start with items that require immediate action, such as disconnected data or concentrated false positives, instead of scanning the entire list from the beginning.</p>
<table><thead><tr><th>Status</th><th>Meaning</th></tr></thead><tbody><tr><td>Review for removal</td><td>No recent alerts or VLM judgments</td></tr><tr><td>Needs confirmation</td><td>Data flow or configuration needs to be checked</td></tr><tr><td>Suspected misapplication</td><td>Accuracy is very low because the scenario does not fit the site</td></tr><tr><td>Performance improvement</td><td>There is room for improvement through tuning</td></tr><tr><td>Feedback needed</td><td>Many alerts, but insufficient review</td></tr><tr><td>Maintain</td><td>Operating normally</td></tr></tbody></table>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="accuracy-analysis---understanding-performance-trends-and-causes">Accuracy Analysis ① — Understanding Performance Trends and Causes<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#accuracy-analysis---understanding-performance-trends-and-causes" class="hash-link" aria-label="Direct link to Accuracy Analysis ① — Understanding Performance Trends and Causes" title="Direct link to Accuracy Analysis ① — Understanding Performance Trends and Causes">​</a></h3>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/image-2-630b6078937a2044b1b3763de1d71965.png" width="100%"></div>
<p>Accuracy Analysis shows alert accuracy and feedback effectiveness together with their calculation logic, and lets operators drill down into detailed metrics by scenario and camera. Because the source alerts and review results behind each metric can be examined, the data provides a basis for decisions such as keeping or removing a scenario.</p>
<p>Change Impact Analysis places scenario edits and model-replacement events on top of the accuracy trend. Operators can therefore understand the relationship between performance changes and configuration changes through recorded history rather than memory or guesswork.</p>
<p>From this screen, operators can confirm:</p>
<ul>
<li>when accuracy changed,</li>
<li>what was changed at that point, and</li>
<li>whether the change actually had an effect.</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="accuracy-analysis---comparing-detection-settings-with-data">Accuracy Analysis ② — Comparing Detection Settings with Data<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#accuracy-analysis---comparing-detection-settings-with-data" class="hash-link" aria-label="Direct link to Accuracy Analysis ② — Comparing Detection Settings with Data" title="Direct link to Accuracy Analysis ② — Comparing Detection Settings with Data">​</a></h3>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/image-3-896e4a0607304eed6c05eb5db53bb31b.png" width="100%"></div>
<p>The objects and thresholds used before VLM judgment also affect the volume and quality of alerts. For example, detecting missing helmets through a <code>person</code> object or a <code>bare head</code> object can produce different results for the same scenario. By comparing the accuracy and alert counts of baseline and alternative settings side by side, operators can use data to decide which camera settings to change first.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="feedback-analysis--checking-the-feedback-loop">Feedback Analysis — Checking the Feedback Loop<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#feedback-analysis--checking-the-feedback-loop" class="hash-link" aria-label="Direct link to Feedback Analysis — Checking the Feedback Loop" title="Direct link to Feedback Analysis — Checking the Feedback Loop">​</a></h3>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/image-4-fea4f3c4cb405498e6110337ebbe77b3.png" width="100%"></div>
<p>EVA suppresses similar false-positive alerts based on user feedback. EVAmentor helps verify whether this filter is working correctly:</p>
<ul>
<li>Did it correctly filter actual false positives?</li>
<li>Did it incorrectly filter real risks?</li>
<li>Which reference images contributed most to the decision?</li>
</ul>
<p>Each event retains its reference image and similarity score, making the reason for filtering traceable. Operators can quickly find incorrectly filtered risks, refine feedback, and check whether a particular reference image is causing excessive filtering.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="ab-testing--improving-without-waiting-for-false-positives-to-recur">A/B Testing — Improving Without Waiting for False Positives to Recur<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#ab-testing--improving-without-waiting-for-false-positives-to-recur" class="hash-link" aria-label="Direct link to A/B Testing — Improving Without Waiting for False Positives to Recur" title="Direct link to A/B Testing — Improving Without Waiting for False Positives to Recur">​</a></h3>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/image-5-3c9d5bb695d78e63dfb96ff8f5a664ad.png" width="100%"></div>
<p>After changing a scenario, operators do not have to wait for the same situation to occur again in the field. They can use previously collected false-positive images to validate a proposed change immediately, even when the false positive occurs only rarely.</p>
<ol>
<li>Re-evaluate past false-positive images with the existing scenario.</li>
<li>Improve the scenario using an AI-generated draft or direct editing.</li>
<li>Apply the revised scenario to the same images and compare the results.</li>
</ol>
<p>Operators can confirm before deployment whether the change reduces false positives while preserving true positives. This moves the process from discovering side effects after deployment to checking them before the change reaches the field.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="image-gallery--turning-review-into-real-improvement">Image Gallery — Turning Review into Real Improvement<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#image-gallery--turning-review-into-real-improvement" class="hash-link" aria-label="Direct link to Image Gallery — Turning Review into Real Improvement" title="Direct link to Image Gallery — Turning Review into Real Improvement">​</a></h3>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/image-6-521443161135ddec7c3cb7b3fbc9bd1d.png" width="100%"></div>
<p>The Image Gallery brings alerts and filtered events together for true-positive and false-positive labeling. Unreviewed items can be compared visually and processed individually or in batches using conditions. Labels do not remain only in EVAmentor; they are also reflected in the corresponding EVA App analysis results and become evidence for future detections.</p>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-the-value-evamentor-creates">2. The Value EVAmentor Creates<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#2-the-value-evamentor-creates" class="hash-link" aria-label="Direct link to 2. The Value EVAmentor Creates" title="Direct link to 2. The Value EVAmentor Creates">​</a></h2>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="managing-the-performance-of-hundreds-of-cameras">Managing the Performance of Hundreds of Cameras<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#managing-the-performance-of-hundreds-of-cameras" class="hash-link" aria-label="Direct link to Managing the Performance of Hundreds of Cameras" title="Direct link to Managing the Performance of Hundreds of Cameras">​</a></h3>
<p>Instead of opening every alert one by one, EVAmentor automatically aggregates performance by site, camera, and scenario and presents an action priority. Even as the number of cameras grows, operators can use the same criteria to review the most important problems first. This changes the daily operation from opening alerts at random to working through the highest-priority items in the action list.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="making-invisible-problems-visible">Making Invisible Problems Visible<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#making-invisible-problems-visible" class="hash-link" aria-label="Direct link to Making Invisible Problems Visible" title="Direct link to Making Invisible Problems Visible">​</a></h3>
<p>EVAmentor surfaces incorrect feedback filters, review gaps, disconnected cameras, and scenarios that no longer fit the site through clear statuses and metrics. It also exposes problems that are easy to miss, such as a blind spot where alerts keep arriving but nobody reviews them, or a filter that suppresses real risks along with false positives. Because the basis for each metric is shown, operational decisions can be explained and defended.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="reducing-review-burden-and-building-operational-assets">Reducing Review Burden and Building Operational Assets<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#reducing-review-burden-and-building-operational-assets" class="hash-link" aria-label="Direct link to Reducing Review Burden and Building Operational Assets" title="Direct link to Reducing Review Burden and Building Operational Assets">​</a></h3>
<p>Operators can label images immediately and process selected items in batches. Since the system shows what should be reviewed first, more important events can be handled within the same amount of time. The resulting review history becomes an operational asset for comparing performance before and after a scenario change and tracking long-term trends.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="making-improvement-about-validation-not-guesswork">Making Improvement About Validation, Not Guesswork<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#making-improvement-about-validation-not-guesswork" class="hash-link" aria-label="Direct link to Making Improvement About Validation, Not Guesswork" title="Direct link to Making Improvement About Validation, Not Guesswork">​</a></h3>
<p>Past false-positive images, object and threshold comparisons, AI-generated scenario drafts, and change-impact analysis allow teams to measure the effect from before a change is made through after it is applied. The process shifts from changing settings based on experience and waiting for field alerts to making data-driven changes and validating them immediately.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="human-decisions-make-the-system-smarter-again">Human Decisions Make the System Smarter Again<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#human-decisions-make-the-system-smarter-again" class="hash-link" aria-label="Direct link to Human Decisions Make the System Smarter Again" title="Direct link to Human Decisions Make the System Smarter Again">​</a></h3>
<p>Labels and review results from EVAmentor are reflected in the next decisions made by EVA App. As their effect is measured again, a positive cycle of <strong>measurement → diagnosis → improvement → application → remeasurement</strong> is created. With each iteration, decisions that do not fit the site decrease and the foundation for trusting field alerts grows stronger.</p>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-closing">3. Closing<a href="https://spectrabrain.ai/en/blog_tech/Innovation/evamentor#3-closing" class="hash-link" aria-label="Direct link to 3. Closing" title="Direct link to 3. Closing">​</a></h2>
<p>If EVA is the system that watches over the field, EVAmentor is the solution that <strong>watches over EVA and makes it more accurate</strong>.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Detection results · Feedback collection</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Performance metric aggregation</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Automatic classification of items requiring action</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Evidence image review · True/false-positive confirmation</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">A/B testing · Detection-setting comparison</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    ↓</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Re-measurement of improvement effects</span><br></span></code></pre></div></div>
<p>Detection quality is not created by a single tuning exercise. It is built by operating this loop continuously.</p>]]></content:encoded>
            <category>Tech</category>
            <category>EVA</category>
            <category>AI Agent</category>
            <category>Vision</category>
        </item>
        <item>
            <title><![CDATA[EVA GPU MIG/MPS Optimization Guide]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/mig_mps</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/mig_mps</guid>
            <pubDate>Mon, 22 Jun 2026 14:00:00 GMT</pubDate>
            <description><![CDATA[EVA optimizes not only model inference, but also GPU partitioning and process execution strategies to reliably operate Vision Models and VLMs in large-scale camera environments.]]></description>
            <content:encoded><![CDATA[<p>EVA optimizes not only model inference, but also GPU partitioning and process execution strategies to reliably operate Vision Models and VLMs in large-scale camera environments.
What matters in this process is not simply using a more powerful GPU. When Vision Models and VLMs run together, actual service performance can vary significantly depending on how GPU resources are partitioned and how multiple inference processes are executed.</p>
<p>EVA does not rely solely on model optimization or application-level scheduling.
It determines the optimal configuration for each deployment environment by considering the number of GPUs installed in the server, GPU memory capacity, MIG partitioning availability, MPS effectiveness, and the placement of Vision Workers and vLLM instances.</p>
<p>In other words, EVA is not just a service that runs AI models. It <strong>maximizes system resource efficiency by considering hardware-level GPU configuration and the behavior of the Serving Framework</strong>. This allows EVA to reliably process requests from many cameras, even with limited server resources.</p>
<p>In this article, we compare the effects of MIG and MPS based on actual EVA experiment data from the following three perspectives.</p>
<ul>
<li>
<strong>MIG effectiveness on a multi-GPU server: PRO 5000 x3</strong>
</li>
<li>
<strong>MIG effectiveness on a single-GPU server: PRO 6000 x1</strong>
</li>
<li>
<strong>MPS effectiveness in an environment with many Vision Workers</strong>
</li>
</ul>
<p>Through this analysis, we aim to provide practical criteria for determining which MIG/MPS configuration is most suitable for EVA operation, rather than applying MIG or MPS unconditionally.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/mig_mps-66ea23de2bb68a8464193df60abbc63d.png" width="80%"></div>
<ul>
<li><strong>MIG(Multi-Instance GPU)</strong>: A feature that partitions a single physical GPU into multiple independent GPU instances. It can reduce resource contention by placing Vision and vLLM workloads on separate instances.</li>
<li><strong>MPS(Multi-Process Service)</strong>: A feature that allows multiple CUDA process requests to be coordinated through a single server process. It can reduce context-switching overhead in multi-process environments.</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-eva-inference-architecture">1. EVA Inference Architecture<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#1-eva-inference-architecture" class="hash-link" aria-label="Direct link to 1. EVA Inference Architecture" title="Direct link to 1. EVA Inference Architecture">​</a></h2>
<p>An EVA server runs two major inference layers at the same time.</p>
<ul>
<li><strong>Vision</strong>: Processes various object detection models such as RT-DETRV2, Owl-v2, OmDet, and LLMDet in parallel through multiple Worker processes.</li>
<li><strong>vLLM(VLM Serving)</strong>: Receives Agent requests, decomposes user-defined scenarios into multiple tasks, and determines whether the scenario condition is met through multi-step inference.</li>
</ul>
<p>The key point in EVA’s inference architecture is that Vision and vLLM use the GPU in different ways.</p>
<table><thead><tr><th>Category</th><th>GPU Usage Pattern</th></tr></thead><tbody><tr><td>Vision</td><td>Many Workers frequently generate short inference requests</td></tr><tr><td>vLLM</td><td>Processes relatively large inference workloads through continuous batching</td></tr><tr><td>Mixed execution</td><td>When Vision and vLLM share the same GPU, context-switching overhead can accumulate</td></tr></tbody></table>
<br>
<p>For models used by many cameras, EVA assigns more Workers to those models so that inference requests can be processed in parallel by model type.
On the other hand, vLLM processes multiple requests together through continuous batching, so simply increasing the number of instances does not always lead to higher throughput.</p>
<p>Therefore, EVA determines whether to apply MIG and MPS based on the server configuration and workload characteristics.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-experiment-environment">2. Experiment Environment<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#2-experiment-environment" class="hash-link" aria-label="Direct link to 2. Experiment Environment" title="Direct link to 2. Experiment Environment">​</a></h2>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="21-server-configuration">2.1 Server Configuration<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#21-server-configuration" class="hash-link" aria-label="Direct link to 2.1 Server Configuration" title="Direct link to 2.1 Server Configuration">​</a></h3>
<p>The experiment was conducted using the following two GPU server configurations.</p>
<table><thead><tr><th>Server</th><th>GPU Configuration</th><th>MIG Configuration</th></tr></thead><tbody><tr><td>Server A</td><td>RTX PRO 5000 48GB x 3</td><td>24GB x 6</td></tr><tr><td>Server B</td><td>RTX PRO 6000 96GB x 1</td><td>24GB x 4</td></tr></tbody></table>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="22-service-placement">2.2 Service Placement<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#22-service-placement" class="hash-link" aria-label="Direct link to 2.2 Service Placement" title="Direct link to 2.2 Service Placement">​</a></h3>
<p>Because Vision and vLLM have different GPU usage patterns, it is important to determine which GPU or MIG instance each workload should be placed on.</p>
<p>In general, Vision is configured as the area responsible for object detection requests, while vLLM is configured as the area responsible for Agent VLM-based reasoning requests. When MIG is applied, each remaining GPU resource or MIG slice, excluding the GPU or MIG slice used by Vision, is assigned one vLLM instance.</p>
<p>For example, when MIG is applied on the PRO 5000 x3 server, six 24GB MIG slices are created. One slice is assigned to Vision, and one vLLM instance is placed on each of the remaining five slices. The same approach is used on the PRO 6000 x1 server: among four 24GB MIG slices, one is assigned to Vision and the remaining three are assigned to vLLM instances.</p>
<table><thead><tr><th>Environment</th><th>MIG</th><th>Placement</th><th>Description</th></tr></thead><tbody><tr><td>PRO 5000 x3</td><td>X</td><td>Vision / vLLM / vLLM</td><td>Among three physical GPUs, one is used for Vision and the other two are used for vLLM instances</td></tr><tr><td>PRO 5000 x3</td><td>O</td><td>Vision / vLLM / vLLM / vLLM / vLLM / vLLM</td><td>Among six 24GB MIG slices, one is used for Vision and the remaining five are used for vLLM instances</td></tr><tr><td>PRO 6000 x1</td><td>X</td><td>Vision + vLLM</td><td>Vision and vLLM run together on a single 96GB GPU</td></tr><tr><td>PRO 6000 x1</td><td>O</td><td>Vision / vLLM / vLLM / vLLM</td><td>Among four 24GB MIG slices, one is used for Vision and the remaining three are used for vLLM instances</td></tr></tbody></table>
<br>
<p>This configuration was designed to evaluate whether MIG can improve actual throughput when Vision and vLLM are separated as much as possible and vLLM instances are evenly placed on the remaining GPU resources.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-metric-definitions">3. Metric Definitions<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#3-metric-definitions" class="hash-link" aria-label="Direct link to 3. Metric Definitions" title="Direct link to 3. Metric Definitions">​</a></h2>
<p>In this article, Vision and Agent throughput are compared using the following metrics.</p>
<table><thead><tr><th>Metric</th><th>Definition</th></tr></thead><tbody><tr><td>Vision throughput</td><td><code>req/s</code></td></tr><tr><td>Agent throughput</td><td><code>req/min</code></td></tr></tbody></table>
<br>
<p>The conversion formulas are as follows.</p>
<ul>
<li>Vision <code>req/s</code> = Total Vision requests processed in 1 hour / 3600</li>
<li>Agent <code>req/min</code> = Total VLM responses in 1 hour / 60</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-mig-effect-on-pro-5000-x3">4. MIG Effect on PRO 5000 x3<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#4-mig-effect-on-pro-5000-x3" class="hash-link" aria-label="Direct link to 4. MIG Effect on PRO 5000 x3" title="Direct link to 4. MIG Effect on PRO 5000 x3">​</a></h2>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="41-measurement-results">4.1 Measurement Results<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#41-measurement-results" class="hash-link" aria-label="Direct link to 4.1 Measurement Results" title="Direct link to 4.1 Measurement Results">​</a></h3>
<p>In the PRO 5000 x3 environment, we measured the performance impact of the MIG-enabled configuration under the following conditions.</p>
<ul>
<li>Test scenario: <strong>Single scenario</strong></li>
<li>Vision model configuration: <strong>Single Vision model</strong></li>
</ul>
<p>Therefore, the results below should not be directly compared with the PRO 6000 x1 environment. Instead, they should be interpreted as the performance change <strong>within the same PRO 5000 x3 environment before and after applying the MIG-enabled configuration</strong>.</p>
<table><thead><tr><th>MIG</th><th>MPS</th><th style="text-align:right">VLM responses</th><th style="text-align:right">VLM Latency</th><th style="text-align:right">Vision Throughput (<code>req/s</code>)</th><th style="text-align:right">Agent Throughput (<code>req/min</code>)</th></tr></thead><tbody><tr><td>X</td><td>X</td><td style="text-align:right">2,287</td><td style="text-align:right">10.39 s</td><td style="text-align:right">29.36</td><td style="text-align:right">38.11</td></tr><tr><td>O</td><td>O</td><td style="text-align:right">2,229</td><td style="text-align:right">10.84 s</td><td style="text-align:right">22.63</td><td style="text-align:right">37.15</td></tr></tbody></table>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="42-interpretation">4.2 Interpretation<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#42-interpretation" class="hash-link" aria-label="Direct link to 4.2 Interpretation" title="Direct link to 4.2 Interpretation">​</a></h3>
<p>In the multi-GPU PRO 5000 x3 configuration, increasing the number of vLLM instances through MIG did not produce a meaningful improvement in vLLM throughput.</p>
<p>Because vLLM was already efficiently handling concurrent requests through continuous batching, increasing the number of instances did not directly translate into higher throughput. In addition, the reduced available resources per instance caused by MIG partitioning and the overall workload placement changes appear to have contributed to the decrease in Vision throughput.</p>
<p>From an operational perspective, the following factors should also be considered when applying MIG.</p>
<ul>
<li>MIG partitioning policy management</li>
<li>Instance-level monitoring</li>
<li>Reassignment and recovery procedures in case of failure</li>
<li>Workload-specific instance size adjustment</li>
</ul>
<p>Therefore, in a multi-GPU environment such as PRO 5000 x3, it is more appropriate to start without MIG by default and consider MIG only when resource contention between Vision and vLLM is clearly identified.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-mig-effect-on-pro-6000-x1">5. MIG Effect on PRO 6000 x1<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#5-mig-effect-on-pro-6000-x1" class="hash-link" aria-label="Direct link to 5. MIG Effect on PRO 6000 x1" title="Direct link to 5. MIG Effect on PRO 6000 x1">​</a></h2>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="51-measurement-results">5.1 Measurement Results<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#51-measurement-results" class="hash-link" aria-label="Direct link to 5.1 Measurement Results" title="Direct link to 5.1 Measurement Results">​</a></h3>
<p>In the PRO 6000 x1 environment, we measured the performance impact of the MIG-enabled configuration under the following conditions.</p>
<ul>
<li>Test scenario: <strong>Two or more complex scenarios</strong></li>
<li>Vision model configuration: <strong>Multiple Vision models</strong></li>
</ul>
<p>Therefore, the results below should not be directly compared with the PRO 5000 x3 environment. Instead, they should be interpreted as the performance change <strong>within the same PRO 6000 x1 environment before and after applying the MIG-enabled configuration</strong>.</p>
<table><thead><tr><th>MIG</th><th>MPS</th><th style="text-align:right">VLM responses</th><th style="text-align:right">VLM Latency</th><th style="text-align:right">Vision Throughput (<code>req/s</code>)</th><th style="text-align:right">Agent Throughput (<code>req/min</code>)</th></tr></thead><tbody><tr><td>X</td><td>X</td><td style="text-align:right">712</td><td style="text-align:right">47.20 s</td><td style="text-align:right">20.33</td><td style="text-align:right">11.87</td></tr><tr><td>O</td><td>O</td><td style="text-align:right">1,032</td><td style="text-align:right">32.10 s</td><td style="text-align:right">26.33</td><td style="text-align:right">17.20</td></tr></tbody></table>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="52-interpretation">5.2 Interpretation<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#52-interpretation" class="hash-link" aria-label="Direct link to 5.2 Interpretation" title="Direct link to 5.2 Interpretation">​</a></h3>
<p>In the single-GPU PRO 6000 x1 configuration, the effect of MIG was clearly observed.</p>
<p>Without MIG, Vision Workers and vLLM share a single physical GPU. In this case, many Vision Workers repeatedly generate short inference requests, while vLLM processes relatively large inference workloads. As a result, GPU ownership can switch frequently between workloads.</p>
<p>By applying MIG, the GPU resources used by Vision and vLLM are isolated into hardware-level independent instances. This reduces resource contention between workloads and allows each inference pipeline to run more stably.</p>
<table><thead><tr><th>Metric</th><th style="text-align:right">MIG Disabled</th><th style="text-align:right">MIG Enabled</th><th style="text-align:right">Change</th></tr></thead><tbody><tr><td>VLM responses</td><td style="text-align:right">712</td><td style="text-align:right">1,032</td><td style="text-align:right">+44.9%</td></tr><tr><td>VLM Latency</td><td style="text-align:right">47.20 s</td><td style="text-align:right">32.10 s</td><td style="text-align:right">-32.0%</td></tr><tr><td>Agent Throughput</td><td style="text-align:right">11.87 req/min</td><td style="text-align:right">17.20 req/min</td><td style="text-align:right">+44.9%</td></tr><tr><td>Vision Throughput</td><td style="text-align:right">20.33 req/s</td><td style="text-align:right">26.33 req/s</td><td style="text-align:right">+29.5%</td></tr></tbody></table>
<br>
<p>These results show that MIG can be an effective option in high-density environments where Vision and vLLM must run together on a single GPU. In particular, the more different the GPU usage patterns of Vision and vLLM are, the greater the benefit of hardware-level isolation through MIG can be.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="6-mps-effect-in-an-environment-with-many-vision-workers">6. MPS Effect in an Environment with Many Vision Workers<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#6-mps-effect-in-an-environment-with-many-vision-workers" class="hash-link" aria-label="Direct link to 6. MPS Effect in an Environment with Many Vision Workers" title="Direct link to 6. MPS Effect in an Environment with Many Vision Workers">​</a></h2>
<p>Vision models run through multiple Worker processes that send requests to the GPU at the same time.
When the number of Workers increases, GPU ownership can switch frequently between processes.</p>
<p>By applying MPS, multiple CUDA process requests can be coordinated through a single MPS server. This can reduce context-switching overhead and improve GPU utilization in multi-process environments.</p>
<p>In this experiment, we compared the effect of MPS on the PRO 5000 environment without applying MIG.</p>
<table><thead><tr><th>MIG</th><th>MPS</th><th style="text-align:right">Total requests</th><th style="text-align:right">Total throughput</th><th style="text-align:right">RT-DETRV2</th><th style="text-align:right">Owl-v2</th><th style="text-align:right">OmDet</th><th style="text-align:right">LLMDet</th></tr></thead><tbody><tr><td>X</td><td>X</td><td style="text-align:right">7,131</td><td style="text-align:right">23.770 req/s</td><td style="text-align:right">4.337 req/s</td><td style="text-align:right">5.597 req/s</td><td style="text-align:right">10.953 req/s</td><td style="text-align:right">2.883 req/s</td></tr><tr><td>X</td><td>O</td><td style="text-align:right">7,794</td><td style="text-align:right">25.980 req/s</td><td style="text-align:right">4.490 req/s</td><td style="text-align:right">7.447 req/s</td><td style="text-align:right">11.120 req/s</td><td style="text-align:right">2.923 req/s</td></tr></tbody></table>
<br>
<p>With MPS enabled, total Vision throughput increased by approximately <strong>9.3%, from 23.770 req/s to 25.980 req/s</strong>.</p>
<p>Among the models, Owl-v2 showed the largest throughput improvement, while RT-DETRV2, OmDet, and LLMDet also showed slight improvements. This indicates that MPS can help coordinate multi-process requests more stably in environments with many Vision Workers.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="7-final-conclusion">7. Final Conclusion<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#7-final-conclusion" class="hash-link" aria-label="Direct link to 7. Final Conclusion" title="Direct link to 7. Final Conclusion">​</a></h2>
<p>This experiment confirmed that MIG and MPS are not features that should be applied uniformly in every environment. Instead, they should be applied selectively depending on the GPU configuration and workload characteristics.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="71-multi-gpu-server-pro-5000-x3">7.1 Multi-GPU Server: PRO 5000 x3<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#71-multi-gpu-server-pro-5000-x3" class="hash-link" aria-label="Direct link to 7.1 Multi-GPU Server: PRO 5000 x3" title="Direct link to 7.1 Multi-GPU Server: PRO 5000 x3">​</a></h3>
<p>In an environment such as PRO 5000 x3, where multiple physical GPUs are available, Vision and vLLM can be separated at the physical GPU level. In this case, additionally applying MIG to increase the number of vLLM instances showed limited throughput improvement.</p>
<ul>
<li>vLLM already handles concurrent requests efficiently through continuous batching</li>
<li>Increasing the number of instances does not directly lead to higher throughput</li>
<li>MIG partitioning can increase operational complexity and resource fragmentation</li>
<li>The default recommendation is to start without MIG</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="72-single-gpu-server-pro-6000-x1">7.2 Single-GPU Server: PRO 6000 x1<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#72-single-gpu-server-pro-6000-x1" class="hash-link" aria-label="Direct link to 7.2 Single-GPU Server: PRO 6000 x1" title="Direct link to 7.2 Single-GPU Server: PRO 6000 x1">​</a></h3>
<p>In an environment such as PRO 6000 x1, where Vision and vLLM must run together on a single physical GPU, MIG showed a significant effect.</p>
<ul>
<li>Vision and vLLM have different GPU usage patterns</li>
<li>Sharing a single GPU can increase context-switching overhead</li>
<li>MIG reduces resource contention between workloads</li>
<li>MIG should be considered first for high-density single-GPU configurations</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="73-environment-with-many-vision-workers">7.3 Environment with Many Vision Workers<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#73-environment-with-many-vision-workers" class="hash-link" aria-label="Direct link to 7.3 Environment with Many Vision Workers" title="Direct link to 7.3 Environment with Many Vision Workers">​</a></h3>
<p>MPS can be effective in environments with many Vision Workers.</p>
<ul>
<li>Many Workers generate GPU requests concurrently</li>
<li>GPU ownership switching between processes can increase overhead</li>
<li>MPS increased total Vision throughput by approximately 9.3%</li>
<li>MPS should be considered as a default option for Vision-heavy servers</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="8-operational-recommendations">8. Operational Recommendations<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#8-operational-recommendations" class="hash-link" aria-label="Direct link to 8. Operational Recommendations" title="Direct link to 8. Operational Recommendations">​</a></h2>
<table><thead><tr><th>Operating Environment</th><th>Recommended Configuration</th></tr></thead><tbody><tr><td>Multi-GPU server focused on vLLM</td><td>Start without MIG and introduce MIG only when a clear bottleneck is identified</td></tr><tr><td>Single-GPU server with mixed Vision/VLM workloads</td><td>Consider MIG first</td></tr><tr><td>Server with many Vision Workers</td><td>Consider applying MPS</td></tr><tr><td>Server where Vision and vLLM can be separated by physical GPU</td><td>Prioritize physical GPU-level separation</td></tr><tr><td>Server with insufficient GPU memory</td><td>Prioritize model placement and Worker count adjustment before MIG</td></tr></tbody></table>
<br>
<p>In summary, EVA does not treat MIG and MPS as simple on/off features.
It selects the optimal configuration for each environment by considering server architecture, Vision/VLM placement, Worker count, vLLM execution behavior, and operational complexity.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="9-references">9. References<a href="https://spectrabrain.ai/en/blog_tech/Research/mig_mps#9-references" class="hash-link" aria-label="Direct link to 9. References" title="Direct link to 9. References">​</a></h2>
<p>The following materials were referenced when organizing the analysis framework for this article.</p>
<ul>
<li>NVIDIA Technical Blog: Getting the Most Out of the NVIDIA A100 GPU with Multi-Instance GPU
<a href="https://developer.nvidia.com/blog/getting-the-most-out-of-the-a100-gpu-with-multi-instance-gpu/" target="_blank" rel="noopener noreferrer">https://developer.nvidia.com/blog/getting-the-most-out-of-the-a100-gpu-with-multi-instance-gpu/</a></li>
<li>NVIDIA Technical Blog: Boost GPU Memory Performance with No Code Changes Using NVIDIA CUDA MPS
<a href="https://developer.nvidia.com/blog/boost-gpu-memory-performance-with-no-code-changes-using-nvidia-cuda-mps/" target="_blank" rel="noopener noreferrer">https://developer.nvidia.com/blog/boost-gpu-memory-performance-with-no-code-changes-using-nvidia-cuda-mps/</a></li>
<li>NVIDIA Documentation: Multi-Process Service (MPS)
<a href="https://docs.nvidia.com/deploy/mps/latest/index.html" target="_blank" rel="noopener noreferrer">https://docs.nvidia.com/deploy/mps/latest/index.html</a></li>
<li>vLLM Official Blog
<a href="https://vllm.ai/blog" target="_blank" rel="noopener noreferrer">https://vllm.ai/blog</a></li>
<li>Anyscale Technical Blog: vLLM Throughput Analysis Based on Continuous Batching
<a href="https://www.anyscale.com/blog/continuous-batching-llm-inference" target="_blank" rel="noopener noreferrer">https://www.anyscale.com/blog/continuous-batching-llm-inference</a></li>
</ul>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>VLM</category>
            <category>GPU</category>
            <category>Serving Framework</category>
        </item>
        <item>
            <title><![CDATA[Infrastructure Optimization for Supporting Large-Scale Camera Environments in EVA]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/optimization</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/optimization</guid>
            <pubDate>Sun, 21 Jun 2026 14:00:00 GMT</pubDate>
            <description><![CDATA[EVA has evolved into an architecture that efficiently utilizes overall server resources, rather than relying solely on GPU performance, to provide AI services for more than 100 cameras on a single server.]]></description>
            <content:encoded><![CDATA[<p>EVA has evolved into an architecture that efficiently utilizes overall server resources, rather than relying solely on GPU performance, to provide AI services for more than 100 cameras on a single server.</p>
<p>In large-scale camera environments, simply using a higher-performance GPU is not enough. The system must maintain stable streaming for all cameras while processing AI inference requests from multiple cameras within limited CPU, memory, network, and GPU resources.</p>
<p>If the AI pipeline becomes biased toward a specific camera or model, some cameras may not be analyzed properly, or the delay between event occurrence and alarm delivery may increase. For this reason, EVA optimizes its infrastructure across the entire pipeline, from video ingestion and AI inference to streaming delivery.</p>
<p>In this article, we introduce the key infrastructure optimization technologies EVA applies to reliably support large-scale camera environments.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-vmvlm-separated-architecture">1. VM·VLM Separated Architecture<a href="https://spectrabrain.ai/en/blog_tech/Research/optimization#1-vmvlm-separated-architecture" class="hash-link" aria-label="Direct link to 1. VM·VLM Separated Architecture" title="Direct link to 1. VM·VLM Separated Architecture">​</a></h2>
<p>EVA does not analyze every video frame with a high-cost VLM. Instead, a VM(Vision Model) first checks whether the target object exists and whether basic conditions are met. Only when further reasoning is required does EVA run the VLM(Vision Language Model).</p>
<p>For example, in a PPE non-compliance detection scenario, frames without a person or frames unrelated to the target condition are filtered out at an early stage. VLM inference is performed only when a person is detected and the system needs to determine whether protective equipment is being worn.</p>
<table><thead><tr><th>Item</th><th>Description</th></tr></thead><tbody><tr><td>Optimization Target</td><td>GPU</td></tr><tr><td>Key Approach</td><td>Use VM-based filtering first and run VLM only when needed</td></tr><tr><td>Effect</td><td>Improved camera capacity on the same server from around 20 cameras to up to 100 cameras</td></tr></tbody></table>
<br>
<p>In the initial architecture, GPU computation increased rapidly because every frame was processed mainly through the VLM. In the current architecture, the VM first filters the targets that require further analysis, significantly reducing the number of VLM calls.</p>
<p>With this structure, EVA improved the number of cameras that can be supported on the same server from around 20 to up to 100. This figure is based on internal validation comparing the initial VLM-centric structure with the current architecture, where VM filtering is applied before invoking the VLM only when necessary.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-task-decomposition-and-gpu-parallel-processing">2. Task Decomposition and GPU Parallel Processing<a href="https://spectrabrain.ai/en/blog_tech/Research/optimization#2-task-decomposition-and-gpu-parallel-processing" class="hash-link" aria-label="Direct link to 2. Task Decomposition and GPU Parallel Processing" title="Direct link to 2. Task Decomposition and GPU Parallel Processing">​</a></h2>
<p>EVA does not process inference requests from multiple cameras as one large task. Instead, it decomposes the scenario reasoning process into smaller tasks such as detection step evaluation, exception checking, image description generation, and vectorization.</p>
<p>These decomposed tasks are distributed and processed in parallel across multiple inference instances. This design prevents any single task from blocking the entire pipeline when requests occur simultaneously from many cameras.</p>
<table><thead><tr><th>Item</th><th>Description</th></tr></thead><tbody><tr><td>Optimization Target</td><td>GPU</td></tr><tr><td>Key Approach</td><td>Decompose scenario reasoning into smaller tasks and process them in parallel</td></tr><tr><td>Effect</td><td>Improved scenario throughput by more than 3x, from 360 to 1,192 scenarios per hour</td></tr></tbody></table>
<br>
<p>The key idea behind this approach is to terminate unnecessary inference early. If an early step determines that the alarm condition is not met, EVA does not proceed with additional VLM calls, image description generation, or vectorization.</p>
<p>As a result, GPU resources can be focused only on tasks that require actual reasoning. In real operating environments, this improved scenario throughput from 360 to 1,192 scenarios per hour, more than a 3x increase.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-inference-optimization-based-on-detection-modes">3. Inference Optimization Based on Detection Modes<a href="https://spectrabrain.ai/en/blog_tech/Research/optimization#3-inference-optimization-based-on-detection-modes" class="hash-link" aria-label="Direct link to 3. Inference Optimization Based on Detection Modes" title="Direct link to 3. Inference Optimization Based on Detection Modes">​</a></h2>
<p>Processing every scenario in the same way increases unnecessary GPU usage. EVA analyzes the characteristics of each user-defined scenario and automatically selects the most appropriate detection method.</p>
<p>Scenarios that only require simple object presence checks are handled mainly by the VM, while scenarios that require complex contextual reasoning are processed using VLM-based inference.</p>
<table><thead><tr><th>Detection Mode</th><th>Main Role</th></tr></thead><tbody><tr><td>Simple Mode</td><td>Detects simple object presence such as people, vehicles, or equipment using the VM</td></tr><tr><td>Default Mode</td><td>Separates various scenarios into detection steps and exception conditions for multi-step reasoning</td></tr><tr><td>PPE Mode</td><td>Precisely checks whether protective equipment is worn at the worker level</td></tr><tr><td>Thinking Mode</td><td>Performs context-based reasoning for complex situations such as fire, falling, or risky behavior</td></tr></tbody></table>
<br>
<p>For example, a scenario such as “Notify me when a person is visible” does not require a high-cost VLM and can be processed with the VM alone. On the other hand, a scenario such as “Notify me when a worker enters a dangerous area without wearing a safety helmet” requires additional reasoning because it must evaluate object presence, PPE status, and spatial context together.</p>
<p>By using only the level of model required for each scenario, EVA reduces GPU usage while maintaining detection performance.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-dynamic-worker-allocation-by-model">4. Dynamic Worker Allocation by Model<a href="https://spectrabrain.ai/en/blog_tech/Research/optimization#4-dynamic-worker-allocation-by-model" class="hash-link" aria-label="Direct link to 4. Dynamic Worker Allocation by Model" title="Direct link to 4. Dynamic Worker Allocation by Model">​</a></h2>
<p>EVA dynamically allocates Workers based on the models in use and the number of cameras assigned to each model.</p>
<p>For example, if many cameras are using OMDet at a certain point in time, EVA assigns more Workers to that model. Models with lower usage are kept with minimal resources. This prevents GPU resources from being overly concentrated on a specific model and helps maintain stable overall GPU utilization.</p>
<table><thead><tr><th>Item</th><th>Description</th></tr></thead><tbody><tr><td>Optimization Target</td><td>GPU, Memory</td></tr><tr><td>Key Approach</td><td>Adjust the number of Workers based on camera count and request volume by model</td></tr><tr><td>Operating Metric</td><td>Maintains frame drop rate within 10% in an internal load test with 100 cameras</td></tr></tbody></table>
<br>
<p>If Workers are allocated statically, a model may occupy resources regardless of actual usage, or a heavily used model may not have enough processing capacity.</p>
<p>EVA adjusts the number of Workers based on request volume by model, keeping the rate of AI inference target frames missed due to processing delay within 10%. This figure is an operating metric managed under an internal load test environment with 100 cameras and simultaneous AI inference requests.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-priority-queue-based-request-scheduling">5. Priority Queue-Based Request Scheduling<a href="https://spectrabrain.ai/en/blog_tech/Research/optimization#5-priority-queue-based-request-scheduling" class="hash-link" aria-label="Direct link to 5. Priority Queue-Based Request Scheduling" title="Direct link to 5. Priority Queue-Based Request Scheduling">​</a></h2>
<p>In large-scale camera environments, certain cameras may temporarily generate a large number of requests. If requests are processed in a simple FIFO order, one camera may occupy excessive system resources, delaying inference for other cameras.</p>
<p>EVA manages inference requests by camera using a priority queue so that all cameras can access the AI pipeline fairly.</p>
<table><thead><tr><th>Scheduling Criteria</th><th>Processing Method</th></tr></thead><tbody><tr><td>Cameras with fewer processed requests</td><td>Processed first</td></tr><tr><td>Requests under the same condition</td><td>Processed in arrival order</td></tr><tr><td>When the queue is full</td><td>Older requests from cameras with excessive accumulated requests are removed first</td></tr><tr><td>Frames that are no longer timely</td><td>Excluded from subsequent inference</td></tr></tbody></table>
<br>
<p>This structure prevents inference for other cameras from being delayed even when requests temporarily spike from a specific camera. It also prevents outdated frames from being processed too late and causing incorrect alarms, helping maintain the stability of the real-time AI pipeline.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="6-dynamic-fps-control-based-on-operating-state">6. Dynamic FPS Control Based on Operating State<a href="https://spectrabrain.ai/en/blog_tech/Research/optimization#6-dynamic-fps-control-based-on-operating-state" class="hash-link" aria-label="Direct link to 6. Dynamic FPS Control Based on Operating State" title="Direct link to 6. Dynamic FPS Control Based on Operating State">​</a></h2>
<p>EVA dynamically adjusts video collection and processing frequency based on each camera’s operating state.</p>
<p>A simple connection state, an AI inference state, and a live streaming state all require different levels of processing. If all cameras are processed at the same FPS, unnecessary CPU, GPU, and network usage increases.</p>
<table><thead><tr><th>Item</th><th>Description</th></tr></thead><tbody><tr><td>Optimization Target</td><td>CPU, GPU, Network</td></tr><tr><td>Key Approach</td><td>Dynamically adjust collection, analysis, and streaming FPS based on camera operating state</td></tr><tr><td>Effect</td><td>Reduced CPU usage by approximately 35% when 50 cameras were connected in Monitoring On state</td></tr></tbody></table>
<br>
<p>For example, in a simple connection state, EVA collects only the minimum number of frames needed to maintain the connection. When AI inference is required, only the frames needed for analysis are selected and processed. Higher-FPS streaming is performed only when a user is watching the live video.</p>
<p>In particular, EVA does not decode every collected frame or use every frame for AI analysis. It selectively extracts only the frames needed for inference, minimizing unnecessary video collection, decoding, preprocessing, and inference operations.</p>
<p>Internal validation showed that, under the same server specifications, CPU usage for 50 cameras in Monitoring On state was reduced from around 100% to around 65%, resulting in an approximately 35% reduction in CPU resource usage.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="7-multi-user-streaming-optimization">7. Multi-User Streaming Optimization<a href="https://spectrabrain.ai/en/blog_tech/Research/optimization#7-multi-user-streaming-optimization" class="hash-link" aria-label="Direct link to 7. Multi-User Streaming Optimization" title="Direct link to 7. Multi-User Streaming Optimization">​</a></h2>
<p>In monitoring environments, multiple users often watch the same camera stream at the same time.</p>
<p>In a conventional structure, as the number of users increases, decoding, preprocessing, and encoding may be repeatedly performed for the same video stream, significantly increasing CPU and memory usage. EVA reduces this overhead by sharing video processing results for the same camera.</p>
<table><thead><tr><th>Item</th><th>Description</th></tr></thead><tbody><tr><td>Optimization Target</td><td>CPU, Memory, Network</td></tr><tr><td>Key Approach</td><td>Decode, preprocess, and encode the same camera stream once, then share it with multiple users</td></tr><tr><td>Effect</td><td>Prevents CPU and memory usage from increasing linearly with the number of users</td></tr></tbody></table>
<br>
<p>EVA performs decoding, preprocessing, and encoding only once for the same camera, then shares the generated stream with multiple users. As a result, even when the number of users increases, the additional cost is mainly limited to network transmission.</p>
<p>This allows EVA to maintain stable streaming quality while minimizing server resource usage, even when many users monitor video streams simultaneously.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="closing">Closing<a href="https://spectrabrain.ai/en/blog_tech/Research/optimization#closing" class="hash-link" aria-label="Direct link to Closing" title="Direct link to Closing">​</a></h2>
<p>EVA’s infrastructure efficiency does not depend solely on single-GPU performance. It is the result of optimizing the entire pipeline, including video ingestion, frame selection, AI inference, request scheduling, and streaming delivery.</p>
<p>EVA combines the following technologies to reliably support more cameras on the same server environment.</p>
<table><thead><tr><th>Optimization Technology</th><th>Main Effect</th></tr></thead><tbody><tr><td>VM·VLM Separation</td><td>Minimizes high-cost VLM calls</td></tr><tr><td>Task Decomposition and Parallel Processing</td><td>Improves scenario throughput by more than 3x</td></tr><tr><td>Detection Mode-Based Optimization</td><td>Uses models according to scenario characteristics</td></tr><tr><td>Dynamic Worker Allocation</td><td>Optimizes GPU and memory usage based on request volume by model</td></tr><tr><td>Priority Queue</td><td>Provides fair inference opportunities across cameras</td></tr><tr><td>Dynamic FPS Control</td><td>Reduces CPU, GPU, and network usage</td></tr><tr><td>Multi-User Streaming Optimization</td><td>Eliminates redundant decoding and encoding</td></tr></tbody></table>
<br>
<p>Ultimately, EVA’s large-scale camera capacity is not achieved by a single optimization technique. It is the result of designing the AI model structure and system infrastructure together to use limited server resources more efficiently.</p>
<p>EVA will continue to enhance its infrastructure optimization technologies to reliably support more cameras, more complex scenarios, and more diverse operating environments.</p>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>Optimization</category>
        </item>
        <item>
            <title><![CDATA[EVA on Rebellions NPU: An Optimization Journey for Physical AI Services]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization</guid>
            <pubDate>Mon, 15 Jun 2026 13:17:00 GMT</pubDate>
            <description><![CDATA[EVA is a Physical AI platform that detects dangerous situations, security events, and worker safety conditions in real time from camera streams.]]></description>
            <content:encoded><![CDATA[<p>EVA is a Physical AI platform that detects dangerous situations, security events, and worker safety conditions in real time from camera streams.
For EVA to operate reliably in real-world sites, it is not enough for the AI model to simply run. The system must process simultaneous requests from multiple cameras and deliver results within a response time users can actually perceive when events occur.</p>
<p>EVA has operated Vision Model, Vision Language Model, and Agent pipelines in GPU-based environments. To secure higher power efficiency and more cost-efficient scalability, we validated and optimized whether EVA could also run at commercial-service quality in a Rebellions NPU environment.</p>
<p>Running EVA on NPU was not just a model-porting task. To operate a GPU-centric AI service pipeline stably on NPU, we had to optimize model compilation, input resolution, parallel processing, and CPU-NPU resource placement together.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="key-optimization-points-validated-for-npu-productionization">Key Optimization Points Validated for NPU Productionization<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization#key-optimization-points-validated-for-npu-productionization" class="hash-link" aria-label="Direct link to Key Optimization Points Validated for NPU Productionization" title="Direct link to Key Optimization Points Validated for NPU Productionization">​</a></h2>
<table><thead><tr><th>Category</th><th>Validation Point</th><th>EVA's Optimization Direction</th></tr></thead><tbody><tr><td>Model compatibility</td><td>GPU-based models need to be transformed for NPU execution structure</td><td>Split model structure, fix input shapes, separate preprocessing/postprocessing</td></tr><tr><td>Input resolution</td><td>A balance is required between inference speed and fine-grained visual accuracy</td><td>Instead of shrinking the full image, crop required regions and reconstruct at target resolution</td></tr><tr><td>Detection quality</td><td>Confirm whether quality is maintained after NPU porting, quantization, and resolution tuning</td><td>Scenario-based performance validation using real operational data</td></tr><tr><td>Inference structure</td><td>As VLM requests increase, Vision Encoder and Decoder load management becomes critical</td><td>Decompose detection into smaller tasks and run only required inference</td></tr><tr><td>Concurrency</td><td>A stable execution architecture is needed for many simultaneous camera requests</td><td>Per-core worker placement on NPU and distributed multi-vLLM instances</td></tr><tr><td>Server resource placement</td><td>Alignment across CPU, memory, and NPU is critical in multi-instance operation</td><td>Apply CPU Pinning and NUMA Alignment</td></tr></tbody></table>
<br>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-optimizing-models-for-npu-execution-structure">1. Optimizing Models for NPU Execution Structure<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization#1-optimizing-models-for-npu-execution-structure" class="hash-link" aria-label="Direct link to 1. Optimizing Models for NPU Execution Structure" title="Direct link to 1. Optimizing Models for NPU Execution Structure">​</a></h2>
<p>The first question in NPU environments was whether each model could be transformed into a structure suitable for NPU execution.
In GPU environments, PyTorch-based models can be run relatively flexibly. In NPU environments, however, model structure and input formats must be explicitly defined through a dedicated compiler.</p>
<p>Because EVA's Vision models handle diverse camera inputs and scenarios, we had to optimize not only model execution itself, but also preprocessing, postprocessing, input shape policy, and coordinate restoration.</p>
<p>To address this, EVA reorganized model structures into NPU-friendly forms, fixed input shapes, and rebuilt image preprocessing and output-to-original-coordinate restoration in line with the EVA pipeline.</p>
<p>This was not merely about placing a model on NPU. It was an optimization effort to provide stable response times and consistent inference outputs in real production environments.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-balancing-input-resolution-and-fine-grained-decision-quality">2. Balancing Input Resolution and Fine-Grained Decision Quality<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization#2-balancing-input-resolution-and-fine-grained-decision-quality" class="hash-link" aria-label="Direct link to 2. Balancing Input Resolution and Fine-Grained Decision Quality" title="Direct link to 2. Balancing Input Resolution and Fine-Grained Decision Quality">​</a></h2>
<p>Raising inference speed on NPU requires input-resolution optimization.
But if full images are uniformly downscaled across all scenarios, performance can degrade in cases that require fine-grained decisions.</p>
<p>This trade-off is especially important in PPE compliance scenarios that require precise recognition of small regions. Helmets, masks, goggles, and gloves occupy small portions of an image, so downscaling the full frame can weaken critical visual details.</p>
<p>EVA did not solve this by simply raising full-image resolution.
Instead, EVA first identifies the region required for decision-making, then crops that region and reconstructs it at the model's required input resolution.</p>
<p>For example, when judging PPE compliance, EVA does not pass the full frame directly to VLM. It crops around the person region and scales that crop to the model input size before inference.</p>
<p>This approach offers several advantages.</p>
<ul>
<li>It reduces inference cost by avoiding high-resolution processing of the full image.</li>
<li>It preserves sufficient resolution in the region the model actually needs to inspect.</li>
<li>It mitigates accuracy loss in small-object or fine-state judgment.</li>
</ul>
<p>In other words, EVA does not simply reduce resolution in NPU environments. It distinguishes between regions the model must inspect and regions it does not.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-detection-quality-validation-with-real-operational-data">3. Detection Quality Validation with Real Operational Data<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization#3-detection-quality-validation-with-real-operational-data" class="hash-link" aria-label="Direct link to 3. Detection Quality Validation with Real Operational Data" title="Direct link to 3. Detection Quality Validation with Real Operational Data">​</a></h2>
<p>For NPU productionization, speed alone is not enough. Detection quality must be validated as well.
Even if a model runs stably on NPU, deployment requires confirming that decision quality is maintained relative to GPU environments.</p>
<p>EVA built its own validation dataset from real detection data collected across operational environments. Rather than relying on public benchmarks, we validated performance against the scenarios EVA actually needs to judge.</p>
<p>The validation set included field scenarios such as the following.</p>
<table><thead><tr><th>Scenario Type</th><th>Examples</th></tr></thead><tbody><tr><td>Safety</td><td>No helmet, no mask, no gloves, fall detection</td></tr><tr><td>Security</td><td>Loitering, fence crossing, access/presence verification</td></tr><tr><td>Equipment/Operations</td><td>Workers around forklifts, equipment contact, load detection</td></tr><tr><td>Disaster/Environment</td><td>Fire, smoke/flame, oil spills on floors</td></tr><tr><td>Vehicle</td><td>Emergency vehicle detection, police vehicle detection</td></tr></tbody></table>
<br>
<p>Using this data, we validated identical scenarios in both GPU and NPU environments.
As a result, we confirmed that NPU could maintain quality at a level similar to GPU across major scenarios.</p>
<p>That said, scenarios requiring fine visual detail for small objects remain sensitive to resolution changes, so combining this with the crop-based input optimization described above is important.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-task-decomposition-to-improve-vlm-call-efficiency">4. Task Decomposition to Improve VLM Call Efficiency<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization#4-task-decomposition-to-improve-vlm-call-efficiency" class="hash-link" aria-label="Direct link to 4. Task Decomposition to Improve VLM Call Efficiency" title="Direct link to 4. Task Decomposition to Improve VLM Call Efficiency">​</a></h2>
<p>In NPU environments, it is often more effective to filter requests that truly require high-cost inference than to process everything as one large inference task. Since VLM combines a Vision Encoder and an LLM Decoder, request volume and execution stages must be managed systematically when many cameras generate requests at once.</p>
<p>EVA does not run VLM on every frame. VM first checks object existence and baseline conditions, and VLM runs only when needed.</p>
<p>For example, in PPE non-compliance detection, frames without a person are filtered out early. Follow-up reasoning runs only when a person region is confirmed. Complex scenarios are also split into smaller tasks such as detection-stage decisions, exception checks, and alert-message generation, rather than handled as one large request.</p>
<p>This lets EVA manage VLM calls efficiently on NPU and concentrate compute resources on requests that actually require semantic reasoning.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-distributed-multi-npu-workers-and-vllm-instances">5. Distributed Multi-NPU Workers and vLLM Instances<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization#5-distributed-multi-npu-workers-and-vllm-instances" class="hash-link" aria-label="Direct link to 5. Distributed Multi-NPU Workers and vLLM Instances" title="Direct link to 5. Distributed Multi-NPU Workers and vLLM Instances">​</a></h2>
<p>To run EVA as a commercial service on NPU, the system must stably handle simultaneous Vision and VLM requests from many cameras. EVA therefore separated Vision and Agent domains, and distributed NPU cores and vLLM instances by role.</p>
<table><thead><tr><th>Domain</th><th>Workload</th><th>NPU Optimization Direction</th></tr></thead><tbody><tr><td>Vision Worker</td><td>Object-detection requests</td><td>Worker placement by NPU core</td></tr><tr><td>vLLM Instance</td><td>VLM-based situation understanding</td><td>Distributed multi-instance configuration</td></tr><tr><td>Vision Encoder</td><td>Image input processing</td><td>Load mitigation through request distribution</td></tr><tr><td>Agent Pipeline</td><td>VLM inference request control</td><td>Forward only required requests to VLM</td></tr></tbody></table>
<br>
<p>In the Vision domain, object-detection workers are mapped per NPU core so many camera requests can run in parallel. Worker counts are tuned by camera volume and per-model request load so NPU cores are used stably.</p>
<p>In the VLM domain, requests are not concentrated into a single vLLM instance. Multiple vLLM instances are distributed to secure concurrent throughput. This is especially important for balancing load in the Vision Encoder path.</p>
<p>In short, EVA's NPU optimization is not just about adding more NPU hardware. It is about placing Vision workers and vLLM instances according to NPU architecture to improve end-to-end inference pipeline stability.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="6-numa-alignment-and-cpu-pinning">6. NUMA Alignment and CPU Pinning<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization#6-numa-alignment-and-cpu-pinning" class="hash-link" aria-label="Direct link to 6. NUMA Alignment and CPU Pinning" title="Direct link to 6. NUMA Alignment and CPU Pinning">​</a></h2>
<p>NUMA stands for Non-Uniform Memory Access, a memory architecture where access latency varies by CPU socket and memory locality.</p>
<p>In NPU optimization, server resource placement is as important as models and pipelines.
In real production environments, multiple vLLM instances run on a single server. Which CPU cores each process uses, which NUMA node those cores belong to, and how physically close they are to the NPU can all affect end-to-end response time.</p>
<p>In multi-instance environments, it is critical to align data movement paths across CPU, memory, and NPU clearly. For this reason, EVA validated a structure that applies CPU Pinning and NUMA Alignment per vLLM instance.</p>
<table><thead><tr><th>Item</th><th>Description</th></tr></thead><tbody><tr><td>CPU Pinning</td><td>Fix CPU cores used by each process</td></tr><tr><td>NUMA Alignment</td><td>Align CPU/memory usage to the NUMA node nearest the NPU</td></tr><tr><td>Container isolation</td><td>Isolate vLLM instances by Docker container or Kubernetes Pod</td></tr><tr><td>Expected effect</td><td>Reduced cross-instance interference and lower data movement cost</td></tr></tbody></table>
<br>
<p>In final production environments, each vLLM should run in an isolated container or pod, with CPU Pinning and NUMA Alignment applied per deployment.</p>
<p>This is not simple server tuning. It is an infrastructure-level optimization required for stable NPU-based AI service operation.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="closing">Closing<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_optimization#closing" class="hash-link" aria-label="Direct link to Closing" title="Direct link to Closing">​</a></h2>
<p>Throughout this optimization journey, EVA did not treat NPU as a simple GPU replacement.
To run a GPU-centric AI service pipeline stably on NPU, we had to jointly optimize model structure, input pipelines, inference stages, worker placement, vLLM instance configuration, and server resource alignment.</p>
<p>The key elements optimized together for stable EVA operation on NPU were:</p>
<ul>
<li>NPU-aligned model structures</li>
<li>Fixed input pipelines</li>
<li>Crop-based input optimization to offset resolution-loss effects</li>
<li>Detection-quality validation based on real operational data</li>
<li>Distributed processing architecture across Vision and VLM</li>
<li>Resource alignment based on CPU Pinning and NUMA Alignment</li>
</ul>
<p>Through this process, EVA secured a structure that maintains performance in major detection scenarios while handling many concurrent camera requests stably in NPU environments.</p>
<p>Going forward, EVA will continue advancing Vision Encoder parallelization, model quantization, multi-NPU scheduling, and container-based resource isolation in Rebellions NPU environments, expanding into a Physical AI platform that runs reliably across diverse hardware infrastructures.</p>]]></content:encoded>
            <category>Tech</category>
            <category>Physical AI</category>
            <category>NPU</category>
            <category>VLM</category>
        </item>
        <item>
            <title><![CDATA[Optimizing Detection Operations with Meta Agent]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/meta_threshold</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/meta_threshold</guid>
            <pubDate>Fri, 29 May 2026 19:00:00 GMT</pubDate>
            <description><![CDATA[New in EVA v3.0: Meta Agent for Easier and More Accurate Detection Operations]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="new-in-eva-v30-meta-agent-for-easier-and-more-accurate-detection-operations">New in EVA v3.0: Meta Agent for Easier and More Accurate Detection Operations<a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#new-in-eva-v30-meta-agent-for-easier-and-more-accurate-detection-operations" class="hash-link" aria-label="Direct link to New in EVA v3.0: Meta Agent for Easier and More Accurate Detection Operations" title="Direct link to New in EVA v3.0: Meta Agent for Easier and More Accurate Detection Operations">​</a></h2>
<p>To operate EVA reliably for a specific purpose, key settings such as object targets, detection sensitivity, and vision models must be continuously tuned to each scenario and camera environment. The core challenge is that it is very difficult for humans to consistently decide "what value should be set now" by combining historical data with the current scene.</p>
<p>To reduce this operational burden and make desired detections easier and more accurate, EVA plans to apply <strong>Meta Agent</strong>. Meta intelligence will continue to evolve, and in <strong>v3.0, object-sensitivity recommendation and vision-model recommendation</strong> are provided first.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-why-meta-agent-is-needed-more-than-ever">1. Why Meta Agent Is Needed More Than Ever<a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#1-why-meta-agent-is-needed-more-than-ever" class="hash-link" aria-label="Direct link to 1. Why Meta Agent Is Needed More Than Ever" title="Direct link to 1. Why Meta Agent Is Needed More Than Ever">​</a></h2>
<p>In small environments, manual camera-by-camera tuning may appear manageable. But once the number of cameras exceeds 100, the operational reality changes completely.</p>
<p>Each camera has different installation height, field of view, lighting, background reflection, and workflow pattern. Even with the same scenario, false-positive/false-negative patterns vary by camera. At that point, per-camera optimization becomes necessary, and that work requires both significant effort and specialized operational know-how.</p>
<p>Typical pain points in real operations:</p>
<ul>
<li>If object definitions are too broad or ambiguous, false positives increase and unnecessary inference grows.</li>
<li>If sensitivity is too low, non-target objects are detected; if too high, required targets are missed.</li>
<li>Model performance varies by scenario (for example, fire/smoke/fall), camera angle, lighting, and background structure.</li>
<li>Even after changing settings, effects are hard to validate immediately, and manual tuning cost rises sharply with camera count.</li>
</ul>
<p>In short, one-time initial setup cannot sustain long-term detection quality. A data-driven auto-optimization layer is required.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-how-meta-agent-operates">2. How Meta Agent Operates<a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#2-how-meta-agent-operates" class="hash-link" aria-label="Direct link to 2. How Meta Agent Operates" title="Direct link to 2. How Meta Agent Operates">​</a></h2>
<p>Meta Agent periodically checks each camera state and operates based on the following signals:</p>
<ul>
<li>Recent detected-alert data</li>
<li>Accumulated user feedback</li>
<li>Accumulated image volume available for recommendation decisions</li>
</ul>
<p>The key point is this: <strong>the more feedback a camera accumulates, the more frequently Meta Agent can analyze that camera, and the more accurate its recommendations become</strong>.</p>
<p>So Meta Agent is not a one-time analyzer. It continuously adjusts recommendation frequency and quality by reflecting both camera-level operating status and feedback accumulation.</p>
<br>
<p>As shown below, operators can view actual detection alerts together with recommendation messages, while EVA provides sensitivity-adjustment recommendations based on accumulated feedback and detection results. For example, when a specific object is repeatedly detected at a low sensitivity level, EVA can recommend adjusting to a sensitivity value that better fits the current scene.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/meta-482e72b92a05237273f67324d10668dc.png" width="50%"></div>
<br>
<p>This process does not end at recommendation alone. By continuously analyzing real detection data from live environments, it also improves EVA's operating strategy and recommendation precision over time.</p>
<p>As camera count grows and more data accumulates, EVA evolves toward better field fitness.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-object-sensitivity-recommendation-logic">3. Object-Sensitivity Recommendation Logic<a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#3-object-sensitivity-recommendation-logic" class="hash-link" aria-label="Direct link to 3. Object-Sensitivity Recommendation Logic" title="Direct link to 3. Object-Sensitivity Recommendation Logic">​</a></h2>
<p>Meta Agent recommendation logic follows this sequence:</p>
<ul>
<li>Run inference on the same object with <strong>five vision models</strong>.</li>
<li>Perform automatic labeling by <strong>ensembling</strong> model outputs.</li>
<li>Verify whether true/false detections can be separated on <strong>20+ images</strong>.</li>
<li>Generate sensitivity/model recommendations only when separability is confirmed.</li>
</ul>
<p>Why ensemble instead of a single model:</p>
<ul>
<li>Objects falsely detected by a vision model are likely to propagate as false positives into VLM-stage decisions.</li>
<li>If early visual recognition is unstable, higher-level reasoning can also be distorted.</li>
<li>Therefore, establishing object validity through cross-model agreement is more stable than trusting one model output.<sup><a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#user-content-fn-1-695e29" id="user-content-fnref-1-695e29" data-footnote-ref="true" aria-describedby="footnote-label">1</a></sup><sup><a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#user-content-fn-2-695e29" id="user-content-fnref-2-695e29" data-footnote-ref="true" aria-describedby="footnote-label">2</a></sup></li>
</ul>
<p>There is also a clear efficiency benefit:</p>
<ul>
<li>VLM inference consumes substantial resources, so large-scale labeling by repeatedly calling VLM is inefficient.</li>
<li>In contrast, multi-vision-model ensembling can validate large samples faster at lower cost, making it more practical for operations.</li>
</ul>
<p>Internal validation showed that multi-model ensemble labeling achieved <strong>about 30% higher labeling accuracy</strong> than single-judgment labeling.</p>
<p>Because all vision models already infer on the same images during this process,</p>
<ul>
<li>appropriate sensitivity can be recommended immediately for the current model, and</li>
<li>if all scenarios detect the same target, model-change recommendation can also be connected.</li>
</ul>
<p>Recommendation message example:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">False positives are occurring for detection target {object}. Try switching to model {A} and adjusting sensitivity to {0.xx}.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">ℹ️ Model-change recommendation is provided only when all scenarios share the same detection target.</span><br></span></code></pre></div></div>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="closing">Closing<a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#closing" class="hash-link" aria-label="Direct link to Closing" title="Direct link to Closing">​</a></h2>
<p>Meta Agent is EVA's operational-optimization layer that shifts sensitivity/model tuning from intuition-driven manual work to data-driven recommendation.</p>
<p>Starting with object-sensitivity and vision-model recommendation in v3.0, Meta intelligence will continue to evolve to improve scenario-level operational quality over time.</p>
<br>
<hr>
<br>
<!-- -->
<section data-footnotes="true" class="footnotes"><h2 class="anchor anchorWithStickyNavbar_LWe7 sr-only" id="footnote-label">Footnotes<a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#footnote-label" class="hash-link" aria-label="Direct link to Footnotes" title="Direct link to Footnotes">​</a></h2>
<ol>
<li id="user-content-fn-1-695e29">
<p>Wang et al., "Evaluating Object Hallucination in Image Captioning" (EMNLP 2018), <a href="https://arxiv.org/abs/1809.02156" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/1809.02156</a> <a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#user-content-fnref-1-695e29" data-footnote-backref="" aria-label="Back to reference 1" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-2-695e29">
<p>Leng et al., "Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding" (CVPR 2024), <a href="https://arxiv.org/abs/2311.16922" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2311.16922</a> <a href="https://spectrabrain.ai/en/blog_tech/Research/meta_threshold#user-content-fnref-2-695e29" data-footnote-backref="" aria-label="Back to reference 2" class="data-footnote-backref">↩</a></p>
</li>
</ol>
</section>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>Vision Model</category>
            <category>AI Agent</category>
        </item>
        <item>
            <title><![CDATA[Thinking Mode for More Accurate Detection of Immediately Discernible Hazards]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/thinking_mode</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/thinking_mode</guid>
            <pubDate>Thu, 28 May 2026 19:00:00 GMT</pubDate>
            <description><![CDATA[New in EVA v3.0: Thinking Mode for More Accurate Detection of Immediately Discernible Hazards]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="new-in-eva-v30-thinking-mode-for-more-accurate-detection-of-immediately-discernible-hazards">New in EVA v3.0: Thinking Mode for More Accurate Detection of Immediately Discernible Hazards<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#new-in-eva-v30-thinking-mode-for-more-accurate-detection-of-immediately-discernible-hazards" class="hash-link" aria-label="Direct link to New in EVA v3.0: Thinking Mode for More Accurate Detection of Immediately Discernible Hazards" title="Direct link to New in EVA v3.0: Thinking Mode for More Accurate Detection of Immediately Discernible Hazards">​</a></h2>
<p>In EVA v3.0, we introduced <strong>Thinking Mode</strong> to reduce false positives in safety incident detection for hazards that can be identified immediately from a single scene. The key idea is that the VLM does not jump to conclusions. Instead, it performs internal reasoning first, checks false-positive risks and logical inconsistencies from multiple angles, and then makes the final hazard decision.</p>
<p>Thinking Mode is especially suitable for <strong>scenarios where the hazard is immediately discernible on screen, but false-positive risk still needs multi-angle verification</strong>.</p>
<p>In addition, VLMs can show a tendency to align with user queries depending on context, and continuously adding exception rules in a base mode can increase latency and operational complexity.<sup><a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#user-content-fn-1-2ee48e" id="user-content-fnref-1-2ee48e" data-footnote-ref="true" aria-describedby="footnote-label">1</a></sup><sup><a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#user-content-fn-2-2ee48e" id="user-content-fnref-2-2ee48e" data-footnote-ref="true" aria-describedby="footnote-label">2</a></sup></p>
<p>As a result, EVA v3.0 uses the model's internal Thinking process to quickly interpret scenes, evaluate false-positive possibility with common-sense reasoning first, and then issue alerts, improving overall detection reliability.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-why-thinking-mode-is-needed-and-how-it-improves-reliability">1. Why Thinking Mode Is Needed, and How It Improves Reliability<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#1-why-thinking-mode-is-needed-and-how-it-improves-reliability" class="hash-link" aria-label="Direct link to 1. Why Thinking Mode Is Needed, and How It Improves Reliability" title="Direct link to 1. Why Thinking Mode Is Needed, and How It Improves Reliability">​</a></h2>
<p>In a simple query-driven baseline, even visually obvious cases such as "is there a fire?" or "did a person fall?" can still generate false positives due to camera angle, reflections, occlusion, or low-light conditions.</p>
<p>On top of that, a rule-patching approach that keeps adding exception rules whenever false positives are found has clear operational limits:</p>
<ul>
<li>Increasing scenario-management complexity as exception rules accumulate</li>
<li>Potentially longer inference paths and higher latency</li>
<li>Reduced operational consistency due to user-by-user rule variations</li>
</ul>
<p>To address this, Thinking Mode is designed to prioritize "review before decision" rather than "quick assertion."</p>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="core-principles">Core Principles<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#core-principles" class="hash-link" aria-label="Direct link to Core Principles" title="Direct link to Core Principles">​</a></h3>
<ul>
<li><strong>Focus on immediately discernible scenes</strong>: Prioritize incidents that can be judged clearly within a single frame.</li>
<li><strong>Multi-angle false-positive review</strong>: Use internal reasoning to check counter-evidence and logical contradictions first.</li>
<li><strong>Final alerts must match incident items</strong>: Trigger alerts only when detected evidence aligns with user-configured incident items.</li>
</ul>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="decision-guidelines">Decision Guidelines<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#decision-guidelines" class="hash-link" aria-label="Direct link to Decision Guidelines" title="Direct link to Decision Guidelines">​</a></h3>
<p>Thinking Mode explicitly defines the following operating policies at the system-prompt level so the model does not drift into overly deep reasoning.</p>
<ul>
<li><strong>Direct visual evidence first</strong>: Use only directly visible evidence from the main scene.</li>
<li><strong>Judge only active incidents</strong>: Evaluate only active incident items from user input, preserving order.</li>
<li><strong>Return False when ambiguous</strong>: If evidence is weak, noisy, occluded, or ambiguous, return False.</li>
<li><strong>Incident-specific precision guards</strong>:<!-- -->
<ul>
<li>Fall/down: Crouching, kneeling, sitting, or perspective ambiguity alone must not trigger True.</li>
<li>Smoke/spark: Steam, dust, reflection, blur, or lighting artifacts alone must not trigger True.</li>
<li>Fire: Bright light or reflection without direct flame/combustion evidence must not trigger True.</li>
</ul>
</li>
<li><strong>Strict output normalization</strong>: Return per-incident outputs only in the required JSON schema for stable post-processing.</li>
</ul>
<p>In short, Thinking Mode is not just a detector. It is an <strong>operational decision mode that combines evidence-centered judgment with false-positive suppression guards</strong>.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-performance-review">2. Performance Review<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#2-performance-review" class="hash-link" aria-label="Direct link to 2. Performance Review" title="Direct link to 2. Performance Review">​</a></h2>
<p>Thinking Mode is evaluated not only by final right/wrong labels, but by how the agent produces intermediate evidence and matches that evidence against scenario rules. As shown below, the model first reasons over scene cues, then determines whether to alert by matching user-requested incidents.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/alert1-f72cf680b882264156d8d34176793e9a.png" width="60%"></div>
<p align="center"><i>Source: National Information Society Agency (NIA), Korea</i></p>
<br>
<p>This case corresponds to fall/down detection. In Thinking Mode, the model reaches <code>alert: true</code> after checking the following core reasoning points.</p>
<ul>
<li><strong>Scene analysis</strong>: Detect a worker in PPE (white protective suit and hard hat) lying horizontally on the industrial floor.</li>
<li><strong>Posture judgment</strong>: Determine that the posture is not crouching/sitting/working posture, but loss of upright posture with body-on-ground evidence.</li>
<li><strong>Fall guard validation</strong>: Verify that the condition of direct visual evidence for collapse/body-on-ground is satisfied.</li>
<li><strong>Ambiguity check</strong>: Confirm low likelihood of benign alternatives (for example stretching or temporary posture change).</li>
<li><strong>Final decision</strong>: Sufficient direct evidence of a fall/down event, therefore <code>alert: true</code>.</li>
</ul>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/alert2-69b863c1a6e86293f179abd8fa8fc813.png" width="60%"></div>
<br>
<p>In contrast, this case appears like a vehicle fire at first glance, but is actually not a fire. Thinking Mode still avoids single-cue decisions and reviews the scene from multiple angles.</p>
<p>Key checkpoints from intermediate reasoning:</p>
<ul>
<li><strong>Scene-context check</strong>: Confirm nighttime parking-lot context with multiple strong light sources (streetlights/headlights).</li>
<li><strong>Re-interpret suspicious cue</strong>: Treat rear red glow as a possible reflection/light artifact instead of immediately labeling it as flame.</li>
<li><strong>Apply fire guards</strong>: Re-check direct flame shape, combustion signals, and smoke evidence.</li>
<li><strong>Final decision</strong>: No clear flame/combustion evidence, therefore <code>alert: false</code>.</li>
</ul>
<p>This shows that even for visually confusing cues such as strong red illumination, Thinking Mode reviews counter-evidence before deciding, reducing false positives and delivering more reliable alerts.</p>
<br>
<p>As a result, the following scenario-level metrics were measured for Thinking Mode:</p>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="falldown">Fall/Down<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#falldown" class="hash-link" aria-label="Direct link to Fall/Down" title="Direct link to Fall/Down">​</a></h4>
<table><thead><tr><th>Mode</th><th style="text-align:right">Accuracy</th><th style="text-align:right">Precision</th><th style="text-align:right">Recall</th><th style="text-align:right">F1 Score</th></tr></thead><tbody><tr><td>Thinking Mode</td><td style="text-align:right">0.8967</td><td style="text-align:right">0.5814</td><td style="text-align:right">0.8621</td><td style="text-align:right">0.6944</td></tr><tr><td>Base Mode</td><td style="text-align:right">0.6854</td><td style="text-align:right">0.2841</td><td style="text-align:right">0.8621</td><td style="text-align:right">0.4274</td></tr></tbody></table>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="smokespark">Smoke/Spark<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#smokespark" class="hash-link" aria-label="Direct link to Smoke/Spark" title="Direct link to Smoke/Spark">​</a></h4>
<table><thead><tr><th>Mode</th><th style="text-align:right">Accuracy</th><th style="text-align:right">Precision</th><th style="text-align:right">Recall</th><th style="text-align:right">F1 Score</th></tr></thead><tbody><tr><td>Thinking Mode</td><td style="text-align:right">0.9041</td><td style="text-align:right">1.0000</td><td style="text-align:right">0.8158</td><td style="text-align:right">0.8986</td></tr><tr><td>Base Mode</td><td style="text-align:right">0.7671</td><td style="text-align:right">0.9565</td><td style="text-align:right">0.5789</td><td style="text-align:right">0.7213</td></tr></tbody></table>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="fire">Fire<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#fire" class="hash-link" aria-label="Direct link to Fire" title="Direct link to Fire">​</a></h4>
<table><thead><tr><th>Mode</th><th style="text-align:right">Accuracy</th><th style="text-align:right">Precision</th><th style="text-align:right">Recall</th><th style="text-align:right">F1 Score</th></tr></thead><tbody><tr><td>Thinking Mode</td><td style="text-align:right">0.9755</td><td style="text-align:right">0.7627</td><td style="text-align:right">0.9000</td><td style="text-align:right">0.8257</td></tr><tr><td>Base Mode</td><td style="text-align:right">0.9328</td><td style="text-align:right">0.4762</td><td style="text-align:right">0.4000</td><td style="text-align:right">0.4348</td></tr></tbody></table>
<br>
<p>In summary, Thinking Mode significantly improves performance for immediately discernible <strong>fall/down, smoke/spark, and fire</strong> scenarios.</p>
<blockquote>
<p>⚠️ Thinking Mode can also perform well in other scenario types, but for cases requiring detailed human-action/gear-state inspection or strict environment-specific rule reasoning, inference time can become very long. We recommend using it selectively for suitable scenarios.</p>
</blockquote>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-usability-expand-coverage-by-editing-incident-items-without-adding-new-scenarios">3. Usability: Expand Coverage by Editing Incident Items, Without Adding New Scenarios<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#3-usability-expand-coverage-by-editing-incident-items-without-adding-new-scenarios" class="hash-link" aria-label="Direct link to 3. Usability: Expand Coverage by Editing Incident Items, Without Adding New Scenarios" title="Direct link to 3. Usability: Expand Coverage by Editing Incident Items, Without Adding New Scenarios">​</a></h2>
<p>In EVA v3.0, Thinking Mode can expand immediately discernible incident detection by editing incident items, without creating a new scenario each time.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">## 💡 Detection with Thinking Mode</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">When an object is detected, the agent reviews false-positive possibility and logical inconsistencies at the configured detection interval,</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">then decides whether a hazardous incident is present. If hazardous, an alert is triggered.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">### Hazard Incidents (up to 3)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Complex or ambiguous incidents may require longer reasoning time.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Please define incidents that can be clearly judged within a single scene.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Fire outbreak</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">**Good examples**</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Fire outbreak</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Smoke outbreak</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Person collapsed on the floor</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">**Not recommended**</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Hazard detection based on specific human actions</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Scenarios requiring composite object-state reasoning</span><br></span></code></pre></div></div>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="closing">Closing<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#closing" class="hash-link" aria-label="Direct link to Closing" title="Direct link to Closing">​</a></h2>
<p>The core of Thinking Mode is "review before decision" rather than "immediate reaction." This helps lower false positives in immediately discernible incident scenarios and provides alert quality that operators can trust in real deployments.</p>
<br>
<hr>
<br>
<!-- -->
<section data-footnotes="true" class="footnotes"><h2 class="anchor anchorWithStickyNavbar_LWe7 sr-only" id="footnote-label">Footnotes<a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#footnote-label" class="hash-link" aria-label="Direct link to Footnotes" title="Direct link to Footnotes">​</a></h2>
<ol>
<li id="user-content-fn-1-2ee48e">
<p>Wang et al., "Evaluating Object Hallucination in Large Vision-Language Models" (EMNLP 2023), <a href="https://arxiv.org/abs/2305.10355" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2305.10355</a> <a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#user-content-fnref-1-2ee48e" data-footnote-backref="" aria-label="Back to reference 1" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-2-2ee48e">
<p>Anthropic, "Towards Understanding Sycophancy in Language Models" (2023), <a href="https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models" target="_blank" rel="noopener noreferrer">https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models</a> <a href="https://spectrabrain.ai/en/blog_tech/Research/thinking_mode#user-content-fnref-2-2ee48e" data-footnote-backref="" aria-label="Back to reference 2" class="data-footnote-backref">↩</a></p>
</li>
</ol>
</section>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>Vision Model</category>
            <category>AI Agent</category>
        </item>
        <item>
            <title><![CDATA[More Accurate PPE Violation Detection with PPE Mode]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/ppe_mode</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/ppe_mode</guid>
            <pubDate>Thu, 28 May 2026 18:00:00 GMT</pubDate>
            <description><![CDATA[New in EVA v3.0: More Accurate PPE Violation Detection with PPE Mode]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="new-in-eva-v30-more-accurate-ppe-violation-detection-with-ppe-mode">New in EVA v3.0: More Accurate PPE Violation Detection with PPE Mode<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#new-in-eva-v30-more-accurate-ppe-violation-detection-with-ppe-mode" class="hash-link" aria-label="Direct link to New in EVA v3.0: More Accurate PPE Violation Detection with PPE Mode" title="Direct link to New in EVA v3.0: More Accurate PPE Violation Detection with PPE Mode">​</a></h2>
<p>In EVA v3.0, we introduced <strong>PPE Mode</strong>, a redesigned VLM inference pipeline to reduce false positives in PPE non-compliance detection. The key change is moving away from "directly answering user queries" toward an agent flow where the model first describes what it observes and then makes decisions based on that evidence.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-why-did-the-previous-approach-produce-many-false-positives">1. Why did the previous approach produce many false positives?<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#1-why-did-the-previous-approach-produce-many-false-positives" class="hash-link" aria-label="Direct link to 1. Why did the previous approach produce many false positives?" title="Direct link to 1. Why did the previous approach produce many false positives?">​</a></h2>
<p>In PPE non-compliance detection, when user queries are highly directive (for example, "find people not wearing helmets"), some VLMs tend to generate <strong>affirmative, query-aligned responses</strong>. In these cases, even weak visual evidence can still lead to a "non-compliance" answer, increasing false positives.</p>
<p>This is closely related to known issues in VLM/LLM systems, including <strong>hallucination</strong> and <strong>sycophancy</strong> (overly aligning with user intent).<sup><a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#user-content-fn-1-3b20b9" id="user-content-fnref-1-3b20b9" data-footnote-ref="true" aria-describedby="footnote-label">1</a></sup><sup><a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#user-content-fn-2-3b20b9" id="user-content-fnref-2-3b20b9" data-footnote-ref="true" aria-describedby="footnote-label">2</a></sup><sup><a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#user-content-fn-3-3b20b9" id="user-content-fnref-3-3b20b9" data-footnote-ref="true" aria-describedby="footnote-label">3</a></sup></p>
<p>By contrast, prompts focused on description (for example, "describe what you see in the image") reduce pressure to agree with user intent and tend to produce more faithful object-state descriptions.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-eva-v30-agent-upgrade-from-answering-the-query-to-describe-then-decide">2. EVA v3.0 Agent Upgrade: From "Answering the Query" to "Describe, Then Decide"<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#2-eva-v30-agent-upgrade-from-answering-the-query-to-describe-then-decide" class="hash-link" aria-label="Direct link to 2. EVA v3.0 Agent Upgrade: From &quot;Answering the Query&quot; to &quot;Describe, Then Decide&quot;" title="Direct link to 2. EVA v3.0 Agent Upgrade: From &quot;Answering the Query&quot; to &quot;Describe, Then Decide&quot;">​</a></h2>
<p>PPE Mode in EVA v3.0 separates detection into a <strong>3-stage pipeline</strong> so that VLM decisions are less directly biased by user intent.</p>
<p>If the previous approach was close to "user query -&gt; immediate decision," PPE Mode separates evidence formation into "target selection -&gt; body-part verification -&gt; state description -&gt; rule matching" to reduce false alarms.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="core-principles">Core Principles<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#core-principles" class="hash-link" aria-label="Direct link to Core Principles" title="Direct link to Core Principles">​</a></h3>
<ul>
<li><strong>Check work context first</strong>: Select candidate workers using person-level boxes, while also incorporating full-image context for scenarios like elevated work.</li>
<li><strong>Separate wearing state from verifiability</strong>: Distinguish "what is worn" from "whether the required body part is actually visible."</li>
<li><strong>Keep final decisions rule-consistent</strong>: Standardize alert criteria by matching required items, synonyms, and required-body-part visibility together.</li>
</ul>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="3-stage-flow">3-Stage Flow<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#3-stage-flow" class="hash-link" aria-label="Direct link to 3-Stage Flow" title="Direct link to 3-Stage Flow">​</a></h3>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="stage-1-candidate-worker-selection-and-required-body-part-definition">Stage 1. Candidate Worker Selection and Required Body-Part Definition<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#stage-1-candidate-worker-selection-and-required-body-part-definition" class="hash-link" aria-label="Direct link to Stage 1. Candidate Worker Selection and Required Body-Part Definition" title="Direct link to Stage 1. Candidate Worker Selection and Required Body-Part Definition">​</a></h4>
<ul>
<li>
<strong>Enrich Required Equipments</strong>
<ul>
<li>Extract synonyms for required detection items.</li>
</ul>
</li>
<li>
<strong>Find Worker</strong>
<ul>
<li>For each person box detected by the object detector, determine whether that person matches the target work situation.</li>
<li>Since some scenarios (for example, elevated work) require full-scene context, the decision also uses the full image.</li>
<li>If multiple candidates exist, they are processed in parallel.</li>
<li>ℹ️ If there are 5 or more people, only the top 5 largest bounding boxes are evaluated.</li>
<li>If at least one relevant worker is found, proceed to Stage 2.</li>
</ul>
</li>
<li>
<strong>Required Body Parts</strong>
<ul>
<li>Derive body parts needed to verify each required item.</li>
<li>Example: mask -&gt; frontal/side face, helmet -&gt; head</li>
</ul>
</li>
</ul>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="stage-2-worn-item-extraction-and-body-part-check">Stage 2. Worn-Item Extraction and Body-Part Check<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#stage-2-worn-item-extraction-and-body-part-check" class="hash-link" aria-label="Direct link to Stage 2. Worn-Item Extraction and Body-Part Check" title="Direct link to Stage 2. Worn-Item Extraction and Body-Part Check">​</a></h4>
<ul>
<li>
<strong>Explain Worker</strong>
<ul>
<li>Extract worn items (helmet, mask, etc.) for workers detected in Stage 1.</li>
<li>If multiple workers are detected, run in parallel.</li>
</ul>
</li>
<li>
<strong>Check Body Parts</strong>
<ul>
<li>Verify whether required body parts are actually visible for workers detected in Stage 1.</li>
<li>If multiple workers are detected, run in parallel.</li>
</ul>
</li>
</ul>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="stage-3-rule-matching-alerting-and-description-generation">Stage 3. Rule Matching, Alerting, and Description Generation<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#stage-3-rule-matching-alerting-and-description-generation" class="hash-link" aria-label="Direct link to Stage 3. Rule Matching, Alerting, and Description Generation" title="Direct link to Stage 3. Rule Matching, Alerting, and Description Generation">​</a></h4>
<ul>
<li>
<strong>Matching Equipments</strong>
<ul>
<li>Trigger an alert when a detected item does not match the required-item synonym set and the required body part is visible.</li>
<li>Include missing required items in the alert message.</li>
<li>Example: PPE non-compliance detection (helmet, mask)</li>
</ul>
</li>
<li>
<strong>Image Description</strong>
<ul>
<li>Generate an image description together with the alert.</li>
</ul>
</li>
</ul>
<p>With this structure, PPE Mode avoids "immediate VLM assertions" and instead makes final decisions through "target selection + body-part check + item extraction + rule matching," improving reproducibility and operational trust.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-performance-review">3. Performance Review<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#3-performance-review" class="hash-link" aria-label="Direct link to 3. Performance Review" title="Direct link to 3. Performance Review">​</a></h2>
<p>PPE Mode is not just evaluated by final right/wrong labels. It operates by <strong>generating intermediate evidence and matching it against scenario rules</strong>. As shown below, it first infers worn items and verifiable body parts, then determines whether to trigger alerts by matching with user scenarios.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/alert-7d5d475b9591b1ebd0662f15407c2a3b.png" width="60%"></div>
<p>For an image like this, the agent generates intermediate outputs such as:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token comment" style="color:hsl(230, 4%, 64%)"># Worn items</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token string" style="color:hsl(119, 34%, 47%)">"worn_items"</span><span class="token plain"> </span><span class="token operator" style="color:hsl(221, 87%, 60%)">=</span><span class="token plain"> </span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">[</span><span class="token string" style="color:hsl(119, 34%, 47%)">"mask"</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"> </span><span class="token string" style="color:hsl(119, 34%, 47%)">"gloves"</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"> </span><span class="token string" style="color:hsl(119, 34%, 47%)">"pants"</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"> </span><span class="token string" style="color:hsl(119, 34%, 47%)">"shoes"</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">]</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token string" style="color:hsl(119, 34%, 47%)">"evidence"</span><span class="token plain"> </span><span class="token operator" style="color:hsl(221, 87%, 60%)">=</span><span class="token plain"> </span><span class="token string" style="color:hsl(119, 34%, 47%)">"The person is wearing a mask, gloves, pants, and shoes."</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token comment" style="color:hsl(230, 4%, 64%)"># Required body-part check</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token string" style="color:hsl(119, 34%, 47%)">"target_region"</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">:</span><span class="token plain"> </span><span class="token string" style="color:hsl(119, 34%, 47%)">"upper head and crown area"</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token string" style="color:hsl(119, 34%, 47%)">"visible"</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">:</span><span class="token plain"> true</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token string" style="color:hsl(119, 34%, 47%)">"evidence"</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">:</span><span class="token plain"> </span><span class="token string" style="color:hsl(119, 34%, 47%)">"The top of the person's head and crown are clearly visible from the front."</span><br></span></code></pre></div></div>
<p>PPE Mode then matches these outputs to the user scenario (for example, missing mask and helmet), and triggers an alert when the condition <strong>required item not detected + required body part visible</strong> is satisfied.
In this example specifically, the person is detected with a mask but without a helmet, and the head region is visible, so a helmet-missing alert is generated.</p>
<br>
<p>As a result, across a <strong>7-scenario PPE dataset</strong>, we observed the following performance gains:</p>
<table><thead><tr><th>Mode</th><th style="text-align:right">Accuracy</th><th style="text-align:right">Precision</th><th style="text-align:right">Recall</th><th style="text-align:right">F1 Score</th></tr></thead><tbody><tr><td>PPE Mode</td><td style="text-align:right">0.7850</td><td style="text-align:right">0.8694</td><td style="text-align:right">0.4010</td><td style="text-align:right">0.5488</td></tr><tr><td>Base Mode</td><td style="text-align:right">0.7116</td><td style="text-align:right">0.6193</td><td style="text-align:right">0.4079</td><td style="text-align:right">0.4918</td></tr></tbody></table>
<br>
<p>Numerically, PPE Mode shows a clear improvement in <strong>detection reliability</strong> over the base mode.</p>
<p>Operationally, the most meaningful change is the <strong>large increase in precision</strong>.
That means a higher portion of "non-compliance" alerts correspond to actual violations, directly reducing unnecessary follow-up checks and alert fatigue.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-usability-change-detection-targets-without-adding-new-scenarios">4. Usability: Change Detection Targets Without Adding New Scenarios<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#4-usability-change-detection-targets-without-adding-new-scenarios" class="hash-link" aria-label="Direct link to 4. Usability: Change Detection Targets Without Adding New Scenarios" title="Direct link to 4. Usability: Change Detection Targets Without Adding New Scenarios">​</a></h2>
<p>PPE Mode generates detection scenarios in the following format, so users can switch required PPE targets without creating new scenarios each time.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">## 💡 PPE mode is enabled for detection.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">When a person is detected, the agent checks whether the person matches the configured work situation at the configured detection interval.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">If any required item is not satisfied, an alert is triggered.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">&gt; ⚠️ For accurate PPE verification, configure detection targets to include the full body in the person box.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">&gt; Example: "person" or "worker"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">&gt; If there are 5 or more people, detection runs only for the top 5 largest boxes.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">### Work Situation (only one can be specified)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Describe the work state or action of the target person.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Seated and performing soldering work</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">### Required Items (up to 3)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Describe the PPE items that must be worn.</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Mask</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">- Helmet</span><br></span></code></pre></div></div>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="closing">Closing<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#closing" class="hash-link" aria-label="Direct link to Closing" title="Direct link to Closing">​</a></h2>
<p>The core of PPE Mode is the shift from "answer first" to "reason from evidence first." By doing so, we focus on reducing false positives in PPE non-compliance detection while improving alert quality that operators can trust in real environments.</p>
<p>As a next step, we will add more real field image cases and publish both quantitative and qualitative analyses on where improvements are most significant.</p>
<br>
<hr>
<br>
<!-- -->
<section data-footnotes="true" class="footnotes"><h2 class="anchor anchorWithStickyNavbar_LWe7 sr-only" id="footnote-label">Footnotes<a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#footnote-label" class="hash-link" aria-label="Direct link to Footnotes" title="Direct link to Footnotes">​</a></h2>
<ol>
<li id="user-content-fn-1-3b20b9">
<p>Wang et al., "Evaluating Object Hallucination in Large Vision-Language Models" (EMNLP 2023), <a href="https://arxiv.org/abs/2305.10355" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2305.10355</a> <a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#user-content-fnref-1-3b20b9" data-footnote-backref="" aria-label="Back to reference 1" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-2-3b20b9">
<p>Bai et al., "A Survey on Hallucination in Large Vision-Language Models" (2024), <a href="https://arxiv.org/abs/2402.00253" target="_blank" rel="noopener noreferrer">https://arxiv.org/abs/2402.00253</a> <a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#user-content-fnref-2-3b20b9" data-footnote-backref="" aria-label="Back to reference 2" class="data-footnote-backref">↩</a></p>
</li>
<li id="user-content-fn-3-3b20b9">
<p>Anthropic, "Towards Understanding Sycophancy in Language Models" (2023), <a href="https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models" target="_blank" rel="noopener noreferrer">https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models</a> <a href="https://spectrabrain.ai/en/blog_tech/Research/ppe_mode#user-content-fnref-3-3b20b9" data-footnote-backref="" aria-label="Back to reference 3" class="data-footnote-backref">↩</a></p>
</li>
</ol>
</section>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>Vision Model</category>
            <category>AI Agent</category>
        </item>
        <item>
            <title><![CDATA[ToaSt: Decoupled Compression for Faster and More Accurate ViTs (ICML 2026)]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/icml2026</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/icml2026</guid>
            <pubDate>Thu, 14 May 2026 18:00:00 GMT</pubDate>
            <description><![CDATA[The ToaSt paper has been accepted to ICML 2026, one of the top conferences in AI.]]></description>
            <content:encoded><![CDATA[<div class="theme-admonition theme-admonition-success admonition_xJq3 alert alert--success"><div class="admonitionHeading_Gvgb"><span class="admonitionIcon_Rf37"><svg viewBox="0 0 12 16"><path fill-rule="evenodd" d="M6.5 0C3.48 0 1 2.19 1 5c0 .92.55 2.25 1 3 1.34 2.25 1.78 2.78 2 4v1h5v-1c.22-1.22.66-1.75 2-4 .45-.75 1-2.08 1-3 0-2.81-2.48-5-5.5-5zm3.64 7.48c-.25.44-.47.8-.67 1.11-.86 1.41-1.25 2.06-1.45 3.23-.02.05-.02.11-.02.17H5c0-.06 0-.13-.02-.17-.2-1.17-.59-1.83-1.45-3.23-.2-.31-.42-.67-.67-1.11C2.44 6.78 2 5.65 2 5c0-2.2 2.02-4 4.5-4 1.22 0 2.36.42 3.22 1.19C10.55 2.94 11 3.94 11 5c0 .66-.44 1.78-.86 2.48zM4 14h5c-.23 1.14-1.3 2-2.5 2s-2.27-.86-2.5-2z"></path></svg></span>success</div><div class="admonitionContent_BuS1"><p><strong>The ToaSt paper has been accepted to ICML 2026, one of the top conferences in AI.</strong></p></div></div>
<p>Vision Transformers (ViTs) power a wide range of tasks, from classification and detection to segmentation and multimodal backbones.
But their high compute cost often becomes a deployment bottleneck.</p>
<p>In this post, we walk through the ICML 2026 ToaSt paper in detail: the motivation, method design, and key experimental results.</p>
<p>The core idea of ToaSt can be summarized in one line:</p>
<ul>
<li>Decouple MHSA and FFN compression, optimize them with different strategies, and improve the accuracy-efficiency trade-off while avoiding cross-layer propagation issues.</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-background-where-vits-become-expensive">1. Background: Where ViTs become expensive<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#1-background-where-vits-become-expensive" class="hash-link" aria-label="Direct link to 1. Background: Where ViTs become expensive" title="Direct link to 1. Background: Where ViTs become expensive">​</a></h2>
<p>ViT compute mainly comes from two sources:</p>
<ul>
<li><strong>Attention</strong>: roughly (O(N^2)) with token length (N)</li>
<li><strong>FFN</strong>: heavy channel-wise computation around hidden dimension (D)</li>
</ul>
<p>As highlighted in the paper, a standard ViT spends about 61% of FLOPs in FFN and around 19% in attention.
This means attention-only acceleration is not enough; FFN redundancy must be addressed directly.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/fig_1_new-b55c66b0fb0bf8e2ae47fddd56fffc3d.png" width="90%"></div>
<p align="center"><i>Figure 1. ToaSt decouples MHSA and FFN compression to avoid harmful cross-layer propagation.</i></p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-limits-of-prior-approaches">2. Limits of prior approaches<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#2-limits-of-prior-approaches" class="hash-link" aria-label="Direct link to 2. Limits of prior approaches" title="Direct link to 2. Limits of prior approaches">​</a></h2>
<p>The paper groups prior ViT acceleration methods into three categories.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="21-structured-weight-pruning">2.1 Structured Weight Pruning<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#21-structured-weight-pruning" class="hash-link" aria-label="Direct link to 2.1 Structured Weight Pruning" title="Direct link to 2.1 Structured Weight Pruning">​</a></h3>
<ul>
<li>Strength: removes heads/channels/blocks in hardware-friendly structured form</li>
<li>Limitation: often requires long retraining to recover accuracy</li>
<li>Practical issue: ViTs are already expensive to train, so long post-pruning fine-tuning is costly</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="22-token-compression--token-merging">2.2 Token Compression / Token Merging<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#22-token-compression--token-merging" class="hash-link" aria-label="Direct link to 2.2 Token Compression / Token Merging" title="Direct link to 2.2 Token Compression / Token Merging">​</a></h3>
<ul>
<li>Strength: directly reduces attention cost by shrinking (N)</li>
<li>Limitation 1: does not directly target dominant FFN (D^2) complexity</li>
<li>Limitation 2: token decisions propagate globally to later layers, making optimization harder</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="23-joint--hybrid-methods">2.3 Joint / Hybrid Methods<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#23-joint--hybrid-methods" class="hash-link" aria-label="Direct link to 2.3 Joint / Hybrid Methods" title="Direct link to 2.3 Joint / Hybrid Methods">​</a></h3>
<ul>
<li>Strength: can optimize multiple axes at once</li>
<li>Limitation: more coupled optimization and higher tuning complexity</li>
<li>Some approaches are also more hardware/kernel dependent in practice</li>
</ul>
<p>ToaSt takes a different route: not one large coupled optimization, but module-wise decoupled compression.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-method-toast-as-layer-independent-compression">3. Method: ToaSt as Layer-Independent Compression<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#3-method-toast-as-layer-independent-compression" class="hash-link" aria-label="Direct link to 3. Method: ToaSt as Layer-Independent Compression" title="Direct link to 3. Method: ToaSt as Layer-Independent Compression">​</a></h2>
<p>ToaSt preserves each block interface ((N \times D)) while compressing internal computation.
This reduces the cascading side effects caused by global structural changes.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="31-mhsa-coupled-structured-pruning">3.1 MHSA: Coupled Structured Pruning<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#31-mhsa-coupled-structured-pruning" class="hash-link" aria-label="Direct link to 3.1 MHSA: Coupled Structured Pruning" title="Direct link to 3.1 MHSA: Coupled Structured Pruning">​</a></h3>
<p>In MHSA, Q/K/V/Proj are mathematically coupled.
ToaSt enforces synchronized index pruning rather than pruning them independently.</p>
<ul>
<li>Q-K synchronized pruning</li>
<li>V-Proj synchronized pruning</li>
<li>Reduce per-head internal dimension (d_k), while preserving the global interface dimension (D)</li>
</ul>
<p>Importance is computed with a geometric-median-based criterion.
Except for small-model settings, the first layer is typically preserved and later layers are pruned aggressively.
The paper reports that aligned (coupled) pruning significantly mitigates accuracy collapse compared with non-aligned pruning at high pruning ratios.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="32-ffn-token-channel-selection-tcs">3.2 FFN: Token Channel Selection (TCS)<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#32-ffn-token-channel-selection-tcs" class="hash-link" aria-label="Direct link to 3.2 FFN: Token Channel Selection (TCS)" title="Direct link to 3.2 FFN: Token Channel Selection (TCS)">​</a></h3>
<p>FFN has a (D \rightarrow 4D \rightarrow D) structure and dominates total FLOPs.
ToaSt introduces training-free dynamic channel selection (TCS) for FFN.</p>
<p>The paper’s FFN analysis reports three consistent patterns:</p>
<ul>
<li>increasing activation sparsity in deeper layers</li>
<li>collapsing effective rank</li>
<li>high (R^2) reconstruction, indicating strong linear channel redundancy</li>
</ul>
<p>Based on this, TCS samples a subset of tokens, estimates channel importance, and keeps only informative channels.
Its importance metric combines global CLS-driven context and patch saliency; for CLS-free architectures (e.g., Swin), patch-only scoring is used.</p>
<p>The pruning policy is asymmetric:</p>
<ul>
<li>FC1 (expansion) is pruned conservatively</li>
<li>FC2 (reduction) is pruned more aggressively in deeper layers (up to 90% in the reported setup)</li>
</ul>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/swin_base_ffn_analysis-6663990c892759bcb13b53e677da2939.png" width="95%"></div>
<p align="center"><i>Figure 2. Swin-Base FFN analysis: redundancy increases in deeper layers.</i></p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/fig_2_new-a92e46b54dfaa53ac88d4743442a9703.png" width="100%"></div>
<p align="center"><i>Figure 3. ToaSt overview: coupled MHSA pruning + FFN TCS.</i></p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-experimental-setup-and-key-results">4. Experimental setup and key results<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#4-experimental-setup-and-key-results" class="hash-link" aria-label="Direct link to 4. Experimental setup and key results" title="Direct link to 4. Experimental setup and key results">​</a></h2>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="41-setup">4.1 Setup<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#41-setup" class="hash-link" aria-label="Direct link to 4.1 Setup" title="Direct link to 4.1 Setup">​</a></h3>
<ul>
<li>Classification: ImageNet-1K</li>
<li>Downstream transfer: COCO 2017 detection (Cascade R-CNN / Mask R-CNN pipelines)</li>
<li>Backbones: 9 models across DeiT (T/S/B), ViT-MAE (B/L/H), Swin (T/S/B)</li>
<li>Metrics: Top-1/Top-5, FLOPs, and throughput/speedup on H100</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="42-imagenet-results">4.2 ImageNet results<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#42-imagenet-results" class="hash-link" aria-label="Direct link to 4.2 ImageNet results" title="Direct link to 4.2 ImageNet results">​</a></h3>
<p>The main message is simultaneous compute reduction and accuracy gain.</p>
<ol>
<li><strong>ViT-MAE-Huge</strong>: Top-1 <strong>88.52%</strong> (vs. 86.88 baseline, <strong>+1.64%p</strong>), <strong>39.4% FLOPs reduction</strong>, <strong>1.59x throughput</strong></li>
<li><strong>DeiT-Small</strong>: Top-1 <strong>83.40%</strong> (vs. 79.82 baseline), <strong>45.7% FLOPs reduction</strong>, <strong>2.07x throughput</strong></li>
<li><strong>Swin-Base</strong>: Top-1 <strong>85.21%</strong> (vs. 83.50 baseline), <strong>42.7% FLOPs reduction</strong>, <strong>1.28x throughput</strong></li>
</ol>
<p>At similar FLOPs budgets, the paper reports multiple cases where ToaSt outperforms token-compression baselines such as ToMe and DiffRate.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="43-downstream-transfer-coco">4.3 Downstream transfer (COCO)<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#43-downstream-transfer-coco" class="hash-link" aria-label="Direct link to 4.3 Downstream transfer (COCO)" title="Direct link to 4.3 Downstream transfer (COCO)">​</a></h3>
<p>Compressed backbones remain competitive when transferred to detection:</p>
<ul>
<li>Swin-Small: <strong>52.2 box mAP</strong> (baseline 51.9)</li>
<li>Swin-Base variants: <strong>52.2 / 51.8 box mAP</strong></li>
</ul>
<p>This suggests ToaSt removes architectural redundancy rather than only overfitting classification behavior.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-what-the-ablations-show">5. What the ablations show<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#5-what-the-ablations-show" class="hash-link" aria-label="Direct link to 5. What the ablations show" title="Direct link to 5. What the ablations show">​</a></h2>
<p>The ablations separate the contribution of each component:</p>
<ul>
<li>MHSA-only: often improves speed but can hurt accuracy</li>
<li>MHSA + TCS (full ToaSt): adds more speed and recovers (or exceeds) baseline accuracy</li>
</ul>
<p>The FC1/FC2 sensitivity analysis also supports the asymmetric pruning policy:</p>
<ul>
<li>FC1 is more sensitive in early layers</li>
<li>FC2 tolerates aggressive pruning in later layers</li>
</ul>
<p>The paper interprets this as a sign that TCS filters redundant channel noise and can behave like implicit regularization.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="6-practical-takeaways">6. Practical takeaways<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#6-practical-takeaways" class="hash-link" aria-label="Direct link to 6. Practical takeaways" title="Direct link to 6. Practical takeaways">​</a></h2>
<p>From an engineering perspective, ToaSt is attractive because:</p>
<ul>
<li>decoupled modules simplify optimization</li>
<li>FFN-focused reduction targets the dominant compute cost</li>
<li>structured outputs are hardware-friendly on commodity GPUs</li>
<li>larger models appear to need fewer recovery epochs after pruning</li>
</ul>
<p>The reported inverse scaling trend in recovery epochs is especially interesting for large foundation backbones.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="7-limitations-and-future-work">7. Limitations and future work<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#7-limitations-and-future-work" class="hash-link" aria-label="Direct link to 7. Limitations and future work" title="Direct link to 7. Limitations and future work">​</a></h2>
<p>The paper explicitly notes one current limitation: layer-wise pruning ratios are manually tuned.
Future directions include:</p>
<ul>
<li>learnable/automatic ratio optimization</li>
<li>extension to VLM settings</li>
<li>combination with quantization</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="conclusion">Conclusion<a href="https://spectrabrain.ai/en/blog_tech/Research/icml2026#conclusion" class="hash-link" aria-label="Direct link to Conclusion" title="Direct link to Conclusion">​</a></h2>
<p>ToaSt addresses two recurring ViT compression pain points at once:</p>
<ul>
<li>token-only compression is insufficient for FFN-dominant compute</li>
<li>globally coupled pruning often increases retraining cost and instability</li>
</ul>
<p>By decoupling MHSA and FFN compression and tailoring each to its own structure, ToaSt achieves a strong and consistent accuracy-efficiency trade-off across model families.</p>
<p>The key message is simple:</p>
<blockquote>
<p>For ViT acceleration, token reduction alone is not enough.
Channel redundancy in FFN must be addressed, and decoupled module-aware design is a practical way to do it.</p>
</blockquote>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>Vision</category>
        </item>
        <item>
            <title><![CDATA[Real-Time Streaming Rendering Optimization - Improving EVA Architecture Based on Canvas, Web Worker, OffscreenCanvas]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/streaming</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/streaming</guid>
            <pubDate>Mon, 06 Apr 2026 10:00:00 GMT</pubDate>
            <description><![CDATA[Hello, I’m Junhyung Yoo, a frontend engineer on the EVA team.]]></description>
            <content:encoded><![CDATA[<p>Hello, I’m Junhyung Yoo, a frontend engineer on the EVA team.</p>
<p>One of the core features of the EVA service is <strong>real-time streaming</strong>, which allows users to monitor video feeds from dozens of cameras simultaneously. As usage expanded from brief checks to long-term on-site monitoring, unexpected performance bottlenecks began to surface.</p>
<blockquote>
<p><strong>"When I leave the screen on for a long time, the browser gradually slows down and eventually the tab crashes."</strong></p>
</blockquote>
<p>To address this issue, we’d like to share our journey of improving the rendering architecture using <strong>Canvas</strong>, <strong>Web Worker</strong>, and <strong>OffscreenCanvas</strong>.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-background-why-did-problems-appear-over-time">1. Background: Why Did Problems Appear Over Time?<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#1-background-why-did-problems-appear-over-time" class="hash-link" aria-label="Direct link to 1. Background: Why Did Problems Appear Over Time?" title="Direct link to 1. Background: Why Did Problems Appear Over Time?">​</a></h2>
<p>In its early days, EVA adopted a very common streaming approach using the <code>&lt;img&gt;</code> tag combined with <code>Blob (Object URL)</code>.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="previous-approach-blob-based-rendering">Previous Approach (Blob-Based Rendering)<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#previous-approach-blob-based-rendering" class="hash-link" aria-label="Direct link to Previous Approach (Blob-Based Rendering)" title="Direct link to Previous Approach (Blob-Based Rendering)">​</a></h3>
<ol>
<li>MJPEG stream data is received from the server in Blob form.</li>
<li>A temporary URL is generated using <code>URL.createObjectURL(blob)</code>.</li>
<li>The URL is assigned to the <code>src</code> of an <code>&lt;img&gt;</code> tag, allowing the browser to render the image.</li>
</ol>
<p>While this implementation was simple, two critical issues emerged in the specialized environment of <b>long-term monitoring</b>.</p>
<ul>
<li><strong>Memory Overhead:</strong> A unique URL string is generated for every frame (around 30 frames per second). Even when calling <code>revokeObjectURL</code>, delays in the browser’s internal image cache and garbage collection (GC) caused memory usage to continuously increase, eventually leading to <strong>Out of Memory (OOM)</strong> errors.</li>
<li><strong>Main Thread Blocking:</strong> Image decoding occurs on the main (UI) thread. When processing high-resolution video, the event loop is delayed, resulting in UI lag such as slow clicks or scrolling—commonly known as <strong>jank</strong>.</li>
</ul>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-network-tab-analysis-understanding-mjpeg">2. Network Tab Analysis: Understanding MJPEG<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#2-network-tab-analysis-understanding-mjpeg" class="hash-link" aria-label="Direct link to 2. Network Tab Analysis: Understanding MJPEG" title="Direct link to 2. Network Tab Analysis: Understanding MJPEG">​</a></h2>
<p>The first step in improving performance was analyzing the <strong>network layer</strong>. MJPEG streaming behaves differently from typical HTTP requests.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="multipartx-mixed-replace">multipart/x-mixed-replace<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#multipartx-mixed-replace" class="hash-link" aria-label="Direct link to multipart/x-mixed-replace" title="Direct link to multipart/x-mixed-replace">​</a></h3>
<p>MJPEG uses the <code>Content-Type: multipart/x-mixed-replace; boundary=...</code> header, which allows the server to continuously push image frames over a single HTTP connection.</p>
<ul>
<li><strong>Network Tab Characteristics:</strong> The request never completes and remains in a <strong>'Pending'</strong> state. The browser keeps the connection open and continuously receives binary data.</li>
<li><strong>Binary Data Structure:</strong> Each frame consists of JPEG binary data (<code>0xFF 0xD8 ... 0xFF 0xD9</code>) separated by a specific <code>boundary</code> string.</li>
</ul>
<p>Because the previous approach converted this massive stream of binary data into Blobs and parsed it on the main thread, browser load increased exponentially as more data accumulated.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-first-optimization-canvas-and-createimagebitmap">3. First Optimization: Canvas and createImageBitmap<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#3-first-optimization-canvas-and-createimagebitmap" class="hash-link" aria-label="Direct link to 3. First Optimization: Canvas and createImageBitmap" title="Direct link to 3. First Optimization: Canvas and createImageBitmap">​</a></h2>
<p>To move away from memory management that relied heavily on the browser’s garbage collector, we introduced the <strong>Canvas API</strong> and adopted an <strong>explicit memory management</strong> approach.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="asynchronous-bitmap-rendering">Asynchronous Bitmap Rendering<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#asynchronous-bitmap-rendering" class="hash-link" aria-label="Direct link to Asynchronous Bitmap Rendering" title="Direct link to Asynchronous Bitmap Rendering">​</a></h3>
<p>The <code>createImageBitmap</code> API allows images to be <strong>decoded asynchronously in the background</strong> before being rendered to the screen.</p>
<div class="language-typescript codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-typescript codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token comment" style="color:hsl(230, 4%, 64%)">// @src/entities/devices/components/stream/MJPEGStream.tsx</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token comment" style="color:hsl(230, 4%, 64%)">// Immediately release memory after drawing the bitmap on the canvas</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token keyword" style="color:hsl(301, 63%, 40%)">const</span><span class="token plain"> bitmap </span><span class="token operator" style="color:hsl(221, 87%, 60%)">=</span><span class="token plain"> </span><span class="token keyword" style="color:hsl(301, 63%, 40%)">await</span><span class="token plain"> </span><span class="token function" style="color:hsl(221, 87%, 60%)">createImageBitmap</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token plain">blob</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">ctx</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">.</span><span class="token function" style="color:hsl(221, 87%, 60%)">drawImage</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token plain">bitmap</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"> </span><span class="token number" style="color:hsl(35, 99%, 36%)">0</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"> </span><span class="token number" style="color:hsl(35, 99%, 36%)">0</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">bitmap</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">.</span><span class="token function" style="color:hsl(221, 87%, 60%)">close</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token plain"> </span><span class="token comment" style="color:hsl(230, 4%, 64%)">// Explicitly release memory</span><br></span></code></pre></div></div>
<p>The key point of this approach is <code>bitmap.close()</code>. By explicitly destroying bitmap resources after use, we were able to keep memory usage stable. In addition, by eliminating <strong>reflow</strong> caused by changing the <code>src</code> of an <code>&lt;img&gt;</code> tag and switching to GPU-accelerated canvas drawing, overall rendering efficiency was significantly improved.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4-second-optimization-separating-computation-with-web-workers">4. Second Optimization: Separating Computation with Web Workers<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#4-second-optimization-separating-computation-with-web-workers" class="hash-link" aria-label="Direct link to 4. Second Optimization: Separating Computation with Web Workers" title="Direct link to 4. Second Optimization: Separating Computation with Web Workers">​</a></h2>
<p>While rendering became lighter, the task of receiving stream data and extracting JPEG frames from binary data (<strong>boundary parsing</strong>) was still handled by the main thread. Performing real-time string searches on millions of bytes per second places a heavy burden on the CPU.</p>
<p>To solve this, we introduced <strong>Web Workers</strong> and applied a clear division of responsibilities:<br>
<strong>"Data processing in the background, rendering on the main thread."</strong></p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="optimizing-data-transfer-transferable-objects">Optimizing Data Transfer (Transferable Objects)<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#optimizing-data-transfer-transferable-objects" class="hash-link" aria-label="Direct link to Optimizing Data Transfer (Transferable Objects)" title="Direct link to Optimizing Data Transfer (Transferable Objects)">​</a></h3>
<p>When sending large images from a worker to the main thread, copying data results in severe performance degradation. We leveraged <strong><code>Transferable Objects</code></strong> to transfer <strong>ownership of memory without copying</strong>, enabling a zero-copy data flow.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-final-optimization-introducing-offscreencanvas">5. Final Optimization: Introducing OffscreenCanvas<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#5-final-optimization-introducing-offscreencanvas" class="hash-link" aria-label="Direct link to 5. Final Optimization: Introducing OffscreenCanvas" title="Direct link to 5. Final Optimization: Introducing OffscreenCanvas">​</a></h2>
<p>Despite these improvements, the final drawing step still occurred on the main thread. The final piece of the puzzle was <strong>OffscreenCanvas</strong>, which allows control of the canvas itself to be transferred to a worker.</p>
<div align="center"><img src="https://t1.kakaocdn.net/kakao_tech/image/2021/06/images/6_offscreencanvas_demo3.gif" width="60%"><em><p>Even when the main thread is blocked (left), image processing running in the worker
continues to update in real time without interruption.
(Source: Kakao Tech Blog)</p></em></div>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="toward-0-rendering-load-on-the-main-thread">Toward 0% Rendering Load on the Main Thread<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#toward-0-rendering-load-on-the-main-thread" class="hash-link" aria-label="Direct link to Toward 0% Rendering Load on the Main Thread" title="Direct link to Toward 0% Rendering Load on the Main Thread">​</a></h3>
<p>After transferring control using <code>transferControlToOffscreen()</code>, rendering is performed entirely inside the worker.</p>
<div class="language-typescript codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-typescript codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token comment" style="color:hsl(230, 4%, 64%)">// @src/entities/devices/components/stream/mjpeg.worker.ts</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token keyword" style="color:hsl(301, 63%, 40%)">const</span><span class="token plain"> bitmap </span><span class="token operator" style="color:hsl(221, 87%, 60%)">=</span><span class="token plain"> </span><span class="token keyword" style="color:hsl(301, 63%, 40%)">await</span><span class="token plain"> </span><span class="token function" style="color:hsl(221, 87%, 60%)">createImageBitmap</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token plain">blob</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token keyword" style="color:hsl(301, 63%, 40%)">if</span><span class="token plain"> </span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token plain">ctx </span><span class="token operator" style="color:hsl(221, 87%, 60%)">&amp;</span><span class="token plain">amp</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token operator" style="color:hsl(221, 87%, 60%)">&amp;</span><span class="token plain">amp</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token plain"> canvas</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token plain"> </span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">{</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  </span><span class="token comment" style="color:hsl(230, 4%, 64%)">// Worker directly draws on the canvas (0% main-thread interference)</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  ctx</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">.</span><span class="token function" style="color:hsl(221, 87%, 60%)">drawImage</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token plain">bitmap</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"> </span><span class="token number" style="color:hsl(35, 99%, 36%)">0</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"> </span><span class="token number" style="color:hsl(35, 99%, 36%)">0</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  </span><span class="token keyword" style="color:hsl(301, 63%, 40%)">if</span><span class="token plain"> </span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token plain">config</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">.</span><span class="token plain">showArea </span><span class="token operator" style="color:hsl(221, 87%, 60%)">&amp;</span><span class="token plain">amp</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token operator" style="color:hsl(221, 87%, 60%)">&amp;</span><span class="token plain">amp</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token plain"> config</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">.</span><span class="token plain">area</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token plain"> </span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">{</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    </span><span class="token function" style="color:hsl(221, 87%, 60%)">drawPolygonArea</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token plain">ctx</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">,</span><span class="token plain"> config</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">.</span><span class="token plain">area</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><span class="token plain"> </span><span class="token comment" style="color:hsl(230, 4%, 64%)">// Area overlay logic also runs in the worker</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  </span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">}</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain"></span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">}</span><span class="token plain"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">bitmap</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">.</span><span class="token function" style="color:hsl(221, 87%, 60%)">close</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">(</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">)</span><span class="token punctuation" style="color:hsl(119, 34%, 47%)">;</span><br></span></code></pre></div></div>
<p>With this architecture, no matter how heavy the workload on the main thread becomes, streaming video continues to play <strong>smoothly and independently</strong> on a separate thread.</p>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="-browser-compatibility-and-automatic-fallback">🌐 Browser Compatibility and Automatic Fallback<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#-browser-compatibility-and-automatic-fallback" class="hash-link" aria-label="Direct link to 🌐 Browser Compatibility and Automatic Fallback" title="Direct link to 🌐 Browser Compatibility and Automatic Fallback">​</a></h3>
<p>While <code>OffscreenCanvas</code> offers powerful capabilities, browser support varies. In the EVA service, browser features are <strong>automatically detected and conditionally handled</strong> based on the user’s environment.</p>
<table><thead><tr><th style="text-align:left">Browser</th><th style="text-align:left">Supported Version</th><th style="text-align:left">Notes</th></tr></thead><tbody><tr><td style="text-align:left"><strong>Chrome</strong></td><td style="text-align:left">69+</td><td style="text-align:left">Primary support</td></tr><tr><td style="text-align:left"><strong>Edge</strong></td><td style="text-align:left">79+</td><td style="text-align:left">Supported from Chromium-based versions</td></tr><tr><td style="text-align:left"><strong>Firefox</strong></td><td style="text-align:left">105+</td><td style="text-align:left">Enabled by default starting from v105</td></tr><tr><td style="text-align:left"><strong>Safari</strong></td><td style="text-align:left">16.4+</td><td style="text-align:left">Latest macOS/iOS recommended</td></tr><tr><td style="text-align:left"><strong>Opera</strong></td><td style="text-align:left">56+</td><td style="text-align:left">-</td></tr></tbody></table>
<p><strong>EVA’s Adaptive Rendering Strategy:</strong></p>
<ul>
<li><strong>Modern browsers:</strong> Enable <code>OffscreenCanvas</code> to keep main-thread load at 0%.</li>
<li><strong>Older browsers (e.g., Safari 15 or below):</strong> Detect feature availability and automatically fall back to the first optimization approach—<strong>main-thread Canvas rendering</strong>.</li>
</ul>
<p>This ensures a seamless streaming experience across all browser environments.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="6-additional-optimizations-buffer-reuse-and-faster-parsing">6. Additional Optimizations: Buffer Reuse and Faster Parsing<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#6-additional-optimizations-buffer-reuse-and-faster-parsing" class="hash-link" aria-label="Direct link to 6. Additional Optimizations: Buffer Reuse and Faster Parsing" title="Direct link to 6. Additional Optimizations: Buffer Reuse and Faster Parsing">​</a></h2>
<p>Performance is determined by details. We applied several additional optimizations within the worker logic.</p>
<ol>
<li><strong>Fixed Buffer Reuse:</strong> Instead of creating new <code>Uint8Array</code> instances each time, we reused fixed-size buffers and managed data using <code>copyWithin</code>. This significantly reduced the frequency of garbage collection (GC).</li>
<li><strong>High-Speed Parsing with <code>indexOf</code>:</strong> Rather than using simple loops to find matching bytes in binary data, we leveraged the built-in <b><code>indexOf</code></b> method to skip unnecessary byte comparisons. Even this simple optimization dramatically reduced frame drops.</li>
</ol>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="7-conclusion-a-more-robust-eva-monitoring-environment">7. Conclusion: A More Robust EVA Monitoring Environment<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#7-conclusion-a-more-robust-eva-monitoring-environment" class="hash-link" aria-label="Direct link to 7. Conclusion: A More Robust EVA Monitoring Environment" title="Direct link to 7. Conclusion: A More Robust EVA Monitoring Environment">​</a></h2>
<p>Through this optimization effort, the EVA service achieved the following results:</p>
<ul>
<li><strong>Memory Stability:</strong> Memory usage remains stable even during long-term operation, eliminating OOM errors.</li>
<li><strong>UI Responsiveness:</strong> UI interactions such as menu navigation and button clicks remain smooth—even during high-resolution streaming—at a near-native app level.</li>
<li><strong>Stable Frame Rates:</strong> By separating threads, consistent frame rates are maintained regardless of network latency or main-thread load.</li>
</ul>
<p>This project reinforced a key principle of frontend performance optimization:<br>
<strong>"How free you keep the browser’s main thread makes all the difference."</strong></p>
<p>Thank you for reading!</p>
<hr>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="technology-summary">Technology Summary<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#technology-summary" class="hash-link" aria-label="Direct link to Technology Summary" title="Direct link to Technology Summary">​</a></h3>
<ul>
<li><strong>Web Workers API:</strong> Execute computations on background threads</li>
<li><strong>OffscreenCanvas:</strong> Rendering independent of the main thread</li>
<li><strong>createImageBitmap:</strong> Asynchronous image decoding with explicit memory management</li>
<li><strong>Transferable Objects:</strong> High-speed data transfer without copy overhead</li>
</ul>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="references">References<a href="https://spectrabrain.ai/en/blog_tech/Research/streaming#references" class="hash-link" aria-label="Direct link to References" title="Direct link to References">​</a></h3>
<ul>
<li><a href="https://t1.kakaocdn.net/kakao_tech" target="_blank" rel="noopener noreferrer">Kakao Tech Blog - OffscreenCanvas</a></li>
<li><a href="https://developer.mozilla.org/en-US/docs/Web/API/Web_Workers_API" target="_blank" rel="noopener noreferrer">MDN - Web Workers</a></li>
<li><a href="https://developer.mozilla.org/en-US/docs/Web/API/OffscreenCanvas" target="_blank" rel="noopener noreferrer">MDN - OffscreenCanvas</a></li>
<li><a href="https://developer.mozilla.org/en-US/docs/Web/API/createImageBitmap" target="_blank" rel="noopener noreferrer">MDN - createImageBitmap</a></li>
<li><a href="https://developer.mozilla.org/en-US/docs/Web/API/Web_Workers_API/Transferable_objects" target="_blank" rel="noopener noreferrer">MDN - Transferable Objects</a></li>
</ul>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>Frontend</category>
        </item>
        <item>
            <title><![CDATA[How EVA Evolved Requirement Management with Agent-Based Workflows]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Innovation/req_process</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Innovation/req_process</guid>
            <pubDate>Mon, 23 Mar 2026 10:00:00 GMT</pubDate>
            <description><![CDATA[Beyond Simple Intake: Agents Handle Requirements]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="beyond-simple-intake-agents-handle-requirements">Beyond Simple Intake: Agents Handle Requirements<a href="https://spectrabrain.ai/en/blog_tech/Innovation/req_process#beyond-simple-intake-agents-handle-requirements" class="hash-link" aria-label="Direct link to Beyond Simple Intake: Agents Handle Requirements" title="Direct link to Beyond Simple Intake: Agents Handle Requirements">​</a></h2>
<p>EVA uses AI Agents as core executors not only in development, but also throughout the requirement management process.</p>
<p>When a problem needs to be solved or a new improvement idea emerges in the field, the process begins simply by sending an email to <code>eva-req@spectrabrain.ai</code>.</p>
<p>What matters here is that writing the email itself does not need to be complicated.</p>
<p>Requesters do not need to spend time formatting their requests.
Instead, they can briefly describe the pain points they experienced and the improvements they expect.</p>
<p>From that point forward, automated Agent workflows take over the refinement, analysis, and prioritization process.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-even-simple-emails-are-structured-by-agents">1) Even Simple Emails Are Structured by Agents<a href="https://spectrabrain.ai/en/blog_tech/Innovation/req_process#1-even-simple-emails-are-structured-by-agents" class="hash-link" aria-label="Direct link to 1) Even Simple Emails Are Structured by Agents" title="Direct link to 1) Even Simple Emails Are Structured by Agents">​</a></h2>
<p>Requesters focus on context rather than format when sending requirement emails.</p>
<p>The Review Agent transforms these free-form inputs into structured requirements that can be directly reviewed by development teams.</p>
<br>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/mail-54c72fb3f27fd64fa77f4963fa7c49a9.png" width="60%"></div>
<br>
<p>During refinement, the Agent supplements missing context and converts ambiguous expressions into actionable language suitable for technical review.</p>
<p>In other words, this step transforms human-friendly input into a system-readable requirement structure.</p>
<br>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/request-a6be104a680c543732d86fe1e8ca5003.png" width="60%"></div>
<br>
<p>At this stage, the requirement is no longer just a note—it becomes an organized unit that supports implementation review and priority assessment.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-analysis-based-on-eva-manuals-and-logic-documentation">2) Analysis Based on EVA Manuals and Logic Documentation<a href="https://spectrabrain.ai/en/blog_tech/Innovation/req_process#2-analysis-based-on-eva-manuals-and-logic-documentation" class="hash-link" aria-label="Direct link to 2) Analysis Based on EVA Manuals and Logic Documentation" title="Direct link to 2) Analysis Based on EVA Manuals and Logic Documentation">​</a></h2>
<p>Once refined, the requirement is analyzed together with EVA’s internal knowledge base.</p>
<p>This includes:</p>
<ul>
<li>User Manuals</li>
<li>Technical and Logic Documentation</li>
</ul>
<p>Using these documents, the Review Agent performs an initial analysis to identify affected areas, possible solution paths, and implementation priorities.</p>
<p>The analysis covers:</p>
<ul>
<li>Impacted features and logic</li>
<li>Whether the issue can be resolved with existing functionality</li>
<li>Whether new implementation is required</li>
<li>Technical risks and expected impact</li>
<li>Development priority</li>
</ul>
<br>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/analysis-7ca422672027603fff270ba82751eaf7.png" width="60%"></div>
<br>
<p>Through this process, a requirement evolves beyond simply describing <em>what the problem is</em>.</p>
<p>It becomes an executable development unit that explains <em>how it can be solved</em> and <em>why it should be addressed now</em>.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-the-agent-loop-continues-after-release">3) The Agent Loop Continues After Release<a href="https://spectrabrain.ai/en/blog_tech/Innovation/req_process#3-the-agent-loop-continues-after-release" class="hash-link" aria-label="Direct link to 3) The Agent Loop Continues After Release" title="Direct link to 3) The Agent Loop Continues After Release">​</a></h2>
<p>The workflow does not stop after implementation.</p>
<p>Once a release is completed, the Release Agent updates the manuals and logic documentation based on the latest changes.</p>
<p>These updated documents then become the knowledge base for future requirement analysis, allowing Agents to make increasingly accurate decisions as the product evolves.</p>
<p>The key point in EVA’s requirement operations is that this process is continuous.</p>
<p>Requirement collection, analysis, development, release, and documentation updates are all connected as a single automated loop without fragmentation.</p>
<br>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/req_process-fd6d7a3477a7340df91b6daac0dbfa5c.png" width="90%"></div>
<br>
<p>Through this loop, Agents continuously learn from signals such as:</p>
<ul>
<li>Which requirements were actually implemented</li>
<li>How they were resolved</li>
<li>Which original expressions were unclear or inaccurate</li>
</ul>
<p>The Review Agent structures inputs, analyzes them, and proposes priorities.</p>
<p>The Release Agent reflects implementation results back into documentation, enabling more accurate analysis for future incoming requirements.</p>
<p>In practice, this creates meaningful improvements such as:</p>
<ul>
<li>Higher requirement analysis accuracy</li>
<li>Automatic detection and cleanup of duplicate requests</li>
<li>Improved documentation quality</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="conclusion-agents-are-not-features-but-operational-processes">Conclusion: Agents Are Not Features, but Operational Processes<a href="https://spectrabrain.ai/en/blog_tech/Innovation/req_process#conclusion-agents-are-not-features-but-operational-processes" class="hash-link" aria-label="Direct link to Conclusion: Agents Are Not Features, but Operational Processes" title="Direct link to Conclusion: Agents Are Not Features, but Operational Processes">​</a></h2>
<p>In EVA, Agents are not just assistive features for isolated tasks.</p>
<p>From the moment a requirement is submitted by email, Agents refine the content, analyze it based on manuals and logic documentation, derive priorities, and update documentation again after release.</p>
<p>In other words, EVA’s requirement management is not a manual process where people repeatedly format requests and organize follow-up work.</p>
<p>It is a workflow where AI Agents manage the entire lifecycle, structure information into maintainable forms, and accumulate knowledge for future decision-making.</p>
<p>As a result, requirements are no longer consumed as one-time requests.</p>
<p>They accumulate as operational data that helps define the product’s evolution more accurately.</p>
<p>EVA uses Agents not simply to automate tasks, but to turn requirement management itself into a continuously learning and improving operational process.</p>]]></content:encoded>
            <category>Tech</category>
            <category>EVA</category>
            <category>Agent</category>
        </item>
        <item>
            <title><![CDATA[Risk Management in Data Centers with EVA]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Innovation/datacenter_risk</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Innovation/datacenter_risk</guid>
            <pubDate>Tue, 17 Mar 2026 16:00:00 GMT</pubDate>
            <description><![CDATA[1. Introduction: Data Center Fires — A “Billion-Dollar” Threat to Business Continuity 🥵]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-introduction-data-center-fires--a-billion-dollar-threat-to-business-continuity-">1. Introduction: Data Center Fires — A “Billion-Dollar” Threat to Business Continuity 🥵<a href="https://spectrabrain.ai/en/blog_tech/Innovation/datacenter_risk#1-introduction-data-center-fires--a-billion-dollar-threat-to-business-continuity-" class="hash-link" aria-label="Direct link to 1. Introduction: Data Center Fires — A “Billion-Dollar” Threat to Business Continuity 🥵" title="Direct link to 1. Introduction: Data Center Fires — A “Billion-Dollar” Threat to Business Continuity 🥵">​</a></h2>
<p>Recent data center fires have gone beyond physical damage, leading to massive financial liabilities due to service disruptions.</p>
<p>👉 <strong>SK C&amp;C Pangyo Data Center Fire (2022)</strong>:
A fire originating in a lithium-ion UPS battery room caused major service outages, including Kakao, with estimated damages reaching trillions of KRW.</p>
<p>👉 <strong>OVHcloud Fire in France (2021)</strong>:
Triggered by UPS power equipment, this incident resulted in approximately €105M in total damages, with €58M covered by insurance—significantly increasing insurer exposure.</p>
<p>These large-scale incidents highlight that modern AI-era data centers carry risks that can no longer be controlled with traditional physical security measures alone.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="-2-data-center-insurance-structure-and-the-surge-in-ai-gpu-center-premiums">😲 2. Data Center Insurance Structure and the Surge in AI GPU Center Premiums<a href="https://spectrabrain.ai/en/blog_tech/Innovation/datacenter_risk#-2-data-center-insurance-structure-and-the-surge-in-ai-gpu-center-premiums" class="hash-link" aria-label="Direct link to 😲 2. Data Center Insurance Structure and the Surge in AI GPU Center Premiums" title="Direct link to 😲 2. Data Center Insurance Structure and the Surge in AI GPU Center Premiums">​</a></h2>
<p>Data center insurance is typically structured as a bundled package including:</p>
<ul>
<li>Property (buildings and servers)</li>
<li>Business Interruption</li>
<li>Cyber Liability</li>
<li>General Liability</li>
</ul>
<p>👉 <strong>Premium Rates Based on Asset Value</strong>
Property insurance premiums usually range from 0.2% to 0.5% of total asset value. However, AI data centers are now facing rapidly increasing premiums due to elevated risk classifications.</p>
<p>👉 <strong>Risks of High-Density Servers</strong>
AI GPU clusters have significantly higher power density compared to traditional servers, directly increasing fire risk.</p>
<table><thead><tr><th>Server Type</th><th>Power Consumption per Rack</th><th>Key Risk Factors</th></tr></thead><tbody><tr><td>Traditional Servers</td><td>5 – 10 kW</td><td>Standard cooling and power management</td></tr><tr><td>High-Performance Computing</td><td>15 – 25 kW</td><td>Increased thermal management requirements</td></tr><tr><td>GPU Clusters</td><td>40 – 120 kW</td><td>Cable overheating, PDU overload, electrical arcs</td></tr></tbody></table>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="-3-key-underwriting-checklist-from-insurers">😎 3. Key Underwriting Checklist from Insurers<a href="https://spectrabrain.ai/en/blog_tech/Innovation/datacenter_risk#-3-key-underwriting-checklist-from-insurers" class="hash-link" aria-label="Direct link to 😎 3. Key Underwriting Checklist from Insurers" title="Direct link to 😎 3. Key Underwriting Checklist from Insurers">​</a></h2>
<p>Global insurance brokers (Aon, Marsh, FM Global, etc.) evaluate risk based not on facility size, but on <strong>technical measures that reduce incident probability</strong>. EVA provides strong advantages across these evaluation criteria.</p>
<p>👉 <strong>Power Infrastructure Risks</strong><br>
Current Status: 40–50% of data center fires originate from electrical systems. Lithium-ion battery thermal runaway is a major driver of rising premiums.<br>
🌈 <strong>EVA’s Role</strong>: Detects minute temperature variations and thermal anomalies at the battery cell level in real time, significantly reducing lithium battery risks.</p>
<br>
<p>👉 <strong>Advanced Fire Detection and Suppression Systems</strong> <br>
Current Status: Early smoke detection and gas-based suppression systems are top underwriting priorities.<br>
🌈 <strong>EVA’s Role</strong>: AI-powered visual intelligence enables instant detection of flames and smoke, reducing detection time to seconds.</p>
<br>
<p>👉 <strong>Operational Risk (Human Error)</strong><br>
Current Status: Insurers heavily assess 24/7 monitoring systems and thermal inspection practices.<br>
🌈 <strong>EVA’s Role</strong>: Transforms manual, human-dependent inspections into automated AI-driven monitoring.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="️-4-economic-impact-of-eva-adoption-a-risk-engineering-approach">❣️ 4. Economic Impact of EVA Adoption: A Risk Engineering Approach<a href="https://spectrabrain.ai/en/blog_tech/Innovation/datacenter_risk#%EF%B8%8F-4-economic-impact-of-eva-adoption-a-risk-engineering-approach" class="hash-link" aria-label="Direct link to ❣️ 4. Economic Impact of EVA Adoption: A Risk Engineering Approach" title="Direct link to ❣️ 4. Economic Impact of EVA Adoption: A Risk Engineering Approach">​</a></h2>
<p>The insurance market is shifting from post-incident compensation to proactive <strong>risk engineering</strong>—reducing the likelihood of future incidents. AI safety solutions like EVA create a win-win structure for both insurers and policyholders.</p>
<p>👉 <strong>Direct Insurance Premium Reduction (15%–30%)</strong>
Well-implemented risk management systems can lead to premium reductions of 15% to 30%. For data centers worth hundreds of billions of KRW, this translates into direct financial benefits that exceed the cost of deploying the solution.</p>
<br>
<p>👉 <strong>Five Core Risk Control Points</strong>
EVA addresses the key underwriting factors that directly impact insurance premiums:</p>
<ul>
<li>Continuous UPS battery monitoring: Early detection of thermal runaway</li>
<li>High power density 대응: Focused monitoring of cable and PDU overheating in AI GPU environments</li>
<li>Intelligent fire detection: Ultra-fast alerts based on visual data</li>
<li>24/7 uninterrupted monitoring: Detection of unsafe behavior and human error</li>
<li>Faster incident response: Reduced time from alert to action</li>
</ul>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="-5-ai-safety-systems-define-data-center-business-continuity">🥰 5. AI Safety Systems Define Data Center Business Continuity<a href="https://spectrabrain.ai/en/blog_tech/Innovation/datacenter_risk#-5-ai-safety-systems-define-data-center-business-continuity" class="hash-link" aria-label="Direct link to 🥰 5. AI Safety Systems Define Data Center Business Continuity" title="Direct link to 🥰 5. AI Safety Systems Define Data Center Business Continuity">​</a></h2>
<p>From a data center operator’s perspective, EVA is not just a “security CCTV system.”</p>
<ul>
<li><strong>Financial Value</strong>: Reduces insurance premiums and prevents large-scale business interruption losses</li>
<li><strong>Operational Value</strong>: Establishes standards for managing power and fire risks in high-density AI environments</li>
<li><strong>Reputational Value</strong>: Strengthens brand trust as a “safe data center” validated by strict insurance assessments</li>
</ul>
<p>Through a virtuous cycle of
<strong>AI Safety System → Risk Reduction → Insurance Discount</strong>,
data centers can achieve both maximum safety and economic efficiency.</p>]]></content:encoded>
            <category>Tech</category>
            <category>EVA</category>
            <category>AI Safety</category>
            <category>Risk Management</category>
        </item>
        <item>
            <title><![CDATA[EVA x Rebellions: Journey of EVA on NPU]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_journey</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_journey</guid>
            <pubDate>Mon, 16 Mar 2026 13:17:00 GMT</pubDate>
            <description><![CDATA[The integration and optimization journey between spectrabrain.ai EVA and Rebellions NPU clearly demonstrates the future direction of next-generation AI infrastructure. Through this project, we verified that NPU-based architectures can address the high cost and power consumption challenges of traditional GPU-centric infrastructures. In particular, in Physical AI environments—where real-time perception and reasoning are critical—we confirmed the potential to achieve both significant TCO (Total Cost of Ownership) reduction and high performance simultaneously.]]></description>
            <content:encoded><![CDATA[<p><strong>The integration and optimization journey between spectrabrain.ai EVA and Rebellions NPU clearly demonstrates the future direction of next-generation AI infrastructure.</strong> Through this project, we verified that NPU-based architectures can address the high cost and power consumption challenges of traditional GPU-centric infrastructures. In particular, in Physical AI environments—where real-time perception and reasoning are critical—we confirmed the potential to achieve both significant TCO (Total Cost of Ownership) reduction and high performance simultaneously.</p>
<p>Today, we would like to share the <strong>porting process of moving GPU-based models to NPUs, along with the technical challenges behind it</strong>, which many people have been curious about.</p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1-the-npu-porting-process-for-gpu-models">1. The NPU Porting Process for GPU Models<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_journey#1-the-npu-porting-process-for-gpu-models" class="hash-link" aria-label="Direct link to 1. The NPU Porting Process for GPU Models" title="Direct link to 1. The NPU Porting Process for GPU Models">​</a></h2>
<p>Since NPUs are designed to accelerate specific types of computations, newly released models cannot be executed immediately without adaptation. To fully utilize the hardware’s capabilities, several essential steps are required.</p>
<ul>
<li>
<p><strong>Model Conversion</strong></p>
<p>The original models developed in PyTorch or TensorFlow must be <strong>converted into an executable format that the NPU can understand</strong>.
Using the <strong>ATOM Compiler</strong> from Rebellions, the model’s computational graph is analyzed and converted into the <strong><code>.rbln</code> executable format</strong> optimized for the NPU architecture.</p>
</li>
<li>
<p><strong>NPU-Optimized Compilation</strong></p>
<p>The model is compiled into a hardware-optimized executable using the compiler in the <strong>Rebellions SDK (RBLN SDK)</strong>.</p>
<ul>
<li><strong>Graph Optimization</strong>: Removes redundant operations and reorganizes the data flow.</li>
<li><strong>Operator Fusion</strong>: Combines multiple small operations into a single large kernel to reduce memory access and execution overhead.</li>
<li><strong>Data Layout Optimization</strong>: Adjusts tensor layouts to match the NPU memory architecture, improving data access efficiency.</li>
</ul>
</li>
<li>
<p><strong>Quantization</strong></p>
<p>Computational precision is adjusted to match the NPU architecture, improving both performance and memory efficiency.
In the case of EVA, we optimized the model to ensure stable performance under an <strong>FP16-based inference environment</strong>.</p>
</li>
<li>
<p><strong>vLLM Integration and Validation</strong></p>
<p>The optimized model is deployed within the <strong>vLLM-RBLN serving framework</strong>. Key metrics such as <strong>TTFT (Time To First Token)</strong> and <strong>throughput</strong> are measured and validated against GPU-based environments.</p>
</li>
</ul>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2-eva-application-optimization-and-technical-challenges">2. EVA Application Optimization and Technical Challenges<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_journey#2-eva-application-optimization-and-technical-challenges" class="hash-link" aria-label="Direct link to 2. EVA Application Optimization and Technical Challenges" title="Direct link to 2. EVA Application Optimization and Technical Challenges">​</a></h2>
<p>After porting the foundation model, the next step is deploying the actual service layer—the <strong>EVA Application</strong>. During this stage, we have been implementing the following optimization roadmap.</p>
<ul>
<li>
<p><strong>EVA Vision Optimization (1:1 Mapping &amp; Batching)</strong></p>
<p>We mapped NPU cores and Vision Workers in a <strong>1:1 configuration</strong>, eliminating context-switching overhead.
In addition, by applying <strong>continuous batching techniques</strong>, we are building a foundation capable of processing data from hundreds of cameras in real time without latency.</p>
</li>
<li>
<p><strong>EVA Agent Optimization (Reducing VLM Load)</strong></p>
<p>The input resolution of the <strong>Vision-Language Model (VLM)</strong> was standardized to <strong>1280×720</strong>, and a <strong>two-stage reasoning architecture</strong> was applied to minimize unnecessary VLM calls.
This immediately reduces the computational load on the <strong>Vision Encoder</strong>, which is one of the most expensive components in the pipeline.</p>
</li>
<li>
<p><strong>System Memory Management and KV Cache Optimization</strong></p>
<p>In collaboration with Rebellions, we analyzed the memory usage patterns of <strong>vLLM-RBLN instances</strong> and improved resource utilization using a <strong>page-based memory management structure</strong>.
This optimization allows the system to process a larger volume of visual data reliably within the same hardware environment.</p>
</li>
<li>
<p><strong>Parallel Processing of the VLM Vision Encoder</strong></p>
<p>We are also improving the parallel execution architecture of the <strong>Vision Encoder</strong>, which accounts for a large portion of the computation in VLM inference.
By optimizing how Vision Encoder operations are distributed across multiple NPU cores, we aim to significantly improve VLM serving throughput.</p>
</li>
</ul>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-conclusion-evolving-from-poc-to-a-production-ready-solution">3. Conclusion: Evolving from PoC to a Production-Ready Solution<a href="https://spectrabrain.ai/en/blog_tech/Innovation/rebellions_journey#3-conclusion-evolving-from-poc-to-a-production-ready-solution" class="hash-link" aria-label="Direct link to 3. Conclusion: Evolving from PoC to a Production-Ready Solution" title="Direct link to 3. Conclusion: Evolving from PoC to a Production-Ready Solution">​</a></h2>
<p>We are continuously addressing technical challenges discovered during stress testing while refining optimizations that maximize hardware utilization. From <strong>parallel processing of the Vision Encoder through close collaboration with Rebellions</strong> to the <strong>development of an intelligent scheduler within the EVA platform</strong>, every step is part of transforming <strong>“EVA on NPU” from a simple proof-of-concept (PoC) into a production-ready solution.</strong></p>
<p>Ultimately, the success of AI services depends on meeting three essential conditions: <strong>economic efficiency, scalability, and service quality.</strong>
EVA will continue to actively adopt the latest NPU technologies and present a <strong>global standard for Physical AI platforms</strong>—delivering the most competitive TCO and outstanding performance for our customers.</p>]]></content:encoded>
            <category>Tech</category>
            <category>Physical AI</category>
            <category>NPU</category>
            <category>VLM</category>
        </item>
        <item>
            <title><![CDATA[Multi-Frame Based VLM Detection: Moving Beyond Single Image Limits to Temporal Context]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/mfm</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/mfm</guid>
            <pubDate>Wed, 11 Mar 2026 22:00:00 GMT</pubDate>
            <description><![CDATA[Is a Single Frame Enough?]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="is-a-single-frame-enough">Is a Single Frame Enough?<a href="https://spectrabrain.ai/en/blog_tech/Research/mfm#is-a-single-frame-enough" class="hash-link" aria-label="Direct link to Is a Single Frame Enough?" title="Direct link to Is a Single Frame Enough?">​</a></h2>
<p>Recently, Vision-Language Models (VLMs) have demonstrated exceptional performance in understanding individual images. Large-scale multimodal models have theoretically expanded the possibilities of multi-frame reasoning by introducing architectures that process multiple images alongside text prompts.</p>
<p>However, real-world industrial detection scenarios are far more complex than controlled research environments. Problems that seem straightforward with a single frame often lead to various false positives and edge cases in production.</p>
<p>Consider a scene where a person is lying on the floor. Looking at that single moment, it is easy to categorize it as a "collapse." But what if the previous frame showed them stretching, or simply changing posture while working?</p>
<p>In nighttime environments, lens flares, light reflections, or glare can mimic the color patterns of fire, leading to false fire detections when based on a single image. When even humans find it difficult to be certain from a single snapshot, providing a model with only one frame inevitably creates structural limitations.</p>
<br>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/figure1-32742b61756205753c5b234cf97ea088.png" width="80%"></div>
<br>
<p>These cases all share a common problem: a "lack of context."</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="time-is-the-most-powerful-context">Time is the Most Powerful Context<a href="https://spectrabrain.ai/en/blog_tech/Research/mfm#time-is-the-most-powerful-context" class="hash-link" aria-label="Direct link to Time is the Most Powerful Context" title="Direct link to Time is the Most Powerful Context">​</a></h2>
<p>Many detection scenarios inherently rely on a temporal flow.</p>
<p>For instance, "loitering" can only be defined by observing a pattern of staying in the same space for a certain period. Similarly, "long-term abandonment" requires the condition that an object remains unchanged for a specific duration after being placed.</p>
<p>Attempting to solve these problems with a single frame is structurally difficult because the focus must be on "change," not just "state."</p>
<p>We have categorized this into three levels of context:</p>
<ul>
<li><strong>Single Image-based Judgment</strong></li>
<li><strong>Short-term Multi-image Contextual Judgment</strong> (Momentary context)</li>
<li><strong>Temporal Judgment</strong> (Involving long-term flow)</li>
</ul>
<p>In actual operating environments, these three levels coexist. Some scenarios are sufficient with a single frame, some require consecutive frames at intervals of a few seconds, and others require tracking a flow over tens of seconds.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="evas-multi-frame-manager">EVA's Multi Frame Manager<a href="https://spectrabrain.ai/en/blog_tech/Research/mfm#evas-multi-frame-manager" class="hash-link" aria-label="Direct link to EVA's Multi Frame Manager" title="Direct link to EVA's Multi Frame Manager">​</a></h2>
<p>In EVA, user-defined scenarios are not treated as simple text conditions. The system analyzes the "level of context" required by each scenario and determines an appropriate frame collection strategy.</p>
<p>For example, "fainting detection" requires multi-images covering a few seconds before and after the event, rather than a single frame. In contrast, "long-term abandonment" requires continuous frame collection over a specific duration based on a sliding window.</p>
<p>The module responsible for this process is the <strong>Multi Frame Manager</strong>. This module dynamically determines the following based on the scenario characteristics:</p>
<ul>
<li>Number of frames required</li>
<li>Collection intervals</li>
<li>Retention time</li>
<li>Event trigger expansion</li>
</ul>
<p>Collected images are not simply listed. They are delivered to the VLM in a clearly sorted chronological order, accompanied by system prompts that guide the model to compare changes between frames.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="multi-image-based-vlm-inference-strategy">Multi-Image Based VLM Inference Strategy<a href="https://spectrabrain.ai/en/blog_tech/Research/mfm#multi-image-based-vlm-inference-strategy" class="hash-link" aria-label="Direct link to Multi-Image Based VLM Inference Strategy" title="Direct link to Multi-Image Based VLM Inference Strategy">​</a></h2>
<p>When multi-frame input is received, the VLM does more than just return independent detection results. In EVA, we designed the inference structure to <strong>interpret multi-images as a continuous temporal context rather than an independent set of images.</strong></p>
<p>To achieve this, frames are delivered to the model using the following strategies:</p>
<ul>
<li><strong>Chronological Frame Alignment</strong>: Constructs time-series data from past to present to understand causality.</li>
<li><strong>Comparative System Prompts</strong>: Uses instructions like "Identify changes compared to the previous frame" to analyze inter-frame correlations.</li>
<li><strong>Temporal Reasoning</strong>: Derives logical conclusions based on state changes over time rather than fragmented snapshot judgments.</li>
</ul>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="case-study-the-power-of-temporal-context-in-reducing-false-positives">Case Study: The Power of Temporal Context in Reducing False Positives<a href="https://spectrabrain.ai/en/blog_tech/Research/mfm#case-study-the-power-of-temporal-context-in-reducing-false-positives" class="hash-link" aria-label="Direct link to Case Study: The Power of Temporal Context in Reducing False Positives" title="Direct link to Case Study: The Power of Temporal Context in Reducing False Positives">​</a></h3>
<p>The following case demonstrates how fragmented information from a single frame is accurately corrected through the "context" of multiple frames.</p>
<br>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/figure2_eng-a28521bd45c7570686ca725344e33f2b.png" width="100%"></div>
<br>
<ul>
<li><strong>Single Image:</strong> A person is stationary in a low, prone position. A VLM looking only at this moment is highly likely to misinterpret the situation as "Collapse."</li>
<li><strong>Multi-Image:</strong> In the subsequent frames, subtle movements are captured—the person moves their arms to operate a phone and tilts their head to look at the screen.</li>
<li><strong>Result:</strong> Through <strong>Temporal Reasoning</strong>, EVA correctly concludes this is <strong>"Sitting and using a phone detected"</strong>.</li>
</ul>
<br>
<p>The core idea is to guide the model to understand the situation by comparing differences between frames, rather than judging each frame individually.</p>
<p>For high-risk detections like fainting, the model undergoes a process of <strong>Progressive Situation Refinement</strong>:</p>
<ol>
<li><strong>Initial State Identification</strong>: Identifying the target object and initial visual features (e.g., prone posture).</li>
<li><strong>Dynamic Change Detection</strong>: Tracking meaningful changes in body angles or voluntary movements compared to previous frames.</li>
<li><strong>Consistency Verification</strong>: Determining if the posture is a forced freeze due to impact or involves intentional actions.</li>
<li><strong>Final Context Determination</strong>: Distinguishing between visual noise with similar patterns and actual events.</li>
</ol>
<p>This <strong>Temporal Reasoning</strong> structure significantly reduces false positives in edge cases that plague single-image systems, providing much more stable results in real-world operations.</p>
<br>
<table><thead><tr><th rowspan="2">Category</th><th colspan="3">Single Image</th><th colspan="3">Multi Image</th></tr><tr><th>Accuracy</th><th>Precision</th><th>Recall</th><th>Accuracy</th><th>Precision</th><th>Recall</th></tr></thead><tbody><tr><td>No PPE</td><td>0.66</td><td>0.87</td><td>0.68</td><td>0.76</td><td>0.87</td><td>0.82</td></tr><tr><td>No Mask (Working)</td><td>0.94</td><td>0.69</td><td>0.54</td><td>0.93</td><td>0.76</td><td>0.52</td></tr><tr><td>Loitering</td><td>0.49</td><td>0.92</td><td>0.33</td><td>0.63</td><td>0.85</td><td>0.64</td></tr><tr><td>Fainting</td><td>0.87</td><td>1.0</td><td>0.36</td><td>0.96</td><td>1.0</td><td>0.82</td></tr></tbody></table>
<br>
<p>Ultimately, EVA’s multi-frame inference structure is not just about increasing the number of input images—it is an approach that directly integrates <strong>temporal change</strong> into the model's reasoning process.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="the-cost-of-multi-frame-computational-overload">The Cost of Multi-Frame: Computational Overload<a href="https://spectrabrain.ai/en/blog_tech/Research/mfm#the-cost-of-multi-frame-computational-overload" class="hash-link" aria-label="Direct link to The Cost of Multi-Frame: Computational Overload" title="Direct link to The Cost of Multi-Frame: Computational Overload">​</a></h2>
<p>Improvements in accuracy come with a price.</p>
<p>While multi-frame reasoning allows for more visual information, it also leads to <strong>increased computational costs</strong>. In multimodal models, image inputs are generally converted into embeddings via a Vision Encoder before being passed to the LLM, a process that is relatively resource-intensive.</p>
<p>Specifically, multi-frame analysis often encounters the following:</p>
<ul>
<li>Identical or very similar images repeating in a sequence.</li>
<li>Multiple requests referencing the same camera frame.</li>
<li>Multiple queries performed on the same set of images.</li>
</ul>
<p>In these cases, if the Vision Encoder processes the same image repeatedly, it creates unnecessary overhead.</p>
<p>In EVA, we developed a structure that maximizes the <strong>Encoder Cache</strong> feature provided by vLLM to solve this. vLLM offers an <strong>Encoder Cache Manager</strong> that allows the system to cache and reuse Vision Encoder results during multimodal processing.</p>
<ul>
<li><a href="https://docs.vllm.ai/en/latest/api/vllm/v1/core/encoder_cache_manager/" target="_blank" rel="noopener noreferrer">https://docs.vllm.ai/en/latest/api/vllm/v1/core/encoder_cache_manager/</a></li>
</ul>
<p>By leveraging this, we can <strong>reuse previously generated encoder embeddings</strong> for identical image inputs, eliminating the need to repeat Vision Encoder operations. EVA applies a request management structure at the <strong>Agent Layer</strong> to effectively utilize this caching.</p>
<br>
<p>The Agent coordinates requests in the following ways:</p>
<ul>
<li>Organizing requests so that identical image inputs can be reused.</li>
<li>Managing requests based on image units to enable cache hits.</li>
<li>Optimizing request flow to prevent redundant encoding.</li>
</ul>
<p>This allows us to minimize Vision Encoder operations and utilize GPU resources more efficiently, even in a multi-frame analysis environment.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="conclusion">Conclusion<a href="https://spectrabrain.ai/en/blog_tech/Research/mfm#conclusion" class="hash-link" aria-label="Direct link to Conclusion" title="Direct link to Conclusion">​</a></h2>
<p>Multi-frame based VLM inference is an approach that significantly improves situational understanding and detection accuracy compared to single-image analysis.</p>
<p>However, as the number of frames increases, the computational load on the Vision Encoder grows significantly. Therefore, it is crucial to design a system that balances <strong>performance gains with computational efficiency and infrastructure costs.</strong></p>
<p>EVA addresses this by actively utilizing vLLM's Encoder Cache and managing requests through the Agent Layer. Through this architecture, we maintain high inference performance while reducing unnecessary computations, continuously improving GPU efficiency and infrastructure operating costs.</p>
<p>This feature is available starting from <strong>EVA v2.6.0.</strong></p>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>VLM</category>
            <category>AI Agent</category>
        </item>
        <item>
            <title><![CDATA[The Future of AI Services Shown by OpenClaw]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Innovation/openclaw</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Innovation/openclaw</guid>
            <pubDate>Tue, 24 Feb 2026 18:00:00 GMT</pubDate>
            <description><![CDATA[Recently, OpenClaw has been generating significant buzz in the AI community. Running in local environments such as a Mac mini, this service interprets a user’s screen in real time and directly controls various applications — signaling an important shift in how we evaluate AI.]]></description>
            <content:encoded><![CDATA[<p>Recently, OpenClaw has been generating significant buzz in the AI community. Running in local environments such as a Mac mini, this service interprets a user’s screen in real time and directly controls various applications — signaling an important shift in how we evaluate AI.</p>
<p>The competitive edge in AI is no longer defined by <strong>“how large or powerful a foundation model is,”</strong> but rather by
<strong>“how effectively that model can perform complex tasks in real-world applications.”</strong></p>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="a-paradigm-shift-from-performance-to-execution">A Paradigm Shift: From Performance to Execution<a href="https://spectrabrain.ai/en/blog_tech/Innovation/openclaw#a-paradigm-shift-from-performance-to-execution" class="hash-link" aria-label="Direct link to A Paradigm Shift: From Performance to Execution" title="Direct link to A Paradigm Shift: From Performance to Execution">​</a></h2>
<ul>
<li>
<p><strong>Old Paradigm: “How intelligent is it?”</strong>
Until now, the AI industry has focused heavily on the scale and performance of foundation models. Large language models such as GPT-4, Claude, and Gemini competed on parameters, dataset size, and benchmark scores. The central question was: <em>“How smart is the AI?”</em></p>
</li>
<li>
<p><strong>New Paradigm: “How much work can it actually perform?”</strong>
OpenClaw introduces a fundamentally different question: <em>“How effectively can the model perform complex tasks in real-world environments?”</em>
AI value is no longer measured by raw intelligence alone, but by its ability to execute within real computing environments.</p>
</li>
</ul>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="the-openclaw-approach">The OpenClaw Approach<a href="https://spectrabrain.ai/en/blog_tech/Innovation/openclaw#the-openclaw-approach" class="hash-link" aria-label="Direct link to The OpenClaw Approach" title="Direct link to The OpenClaw Approach">​</a></h2>
<ul>
<li>
<p><strong>Real-time screen interpretation and contextual awareness</strong>
The ability to interpret a user’s screen in real time shows that AI can move beyond text processing to understand and respond to visual context. This represents a practical implementation of multimodal AI.</p>
</li>
<li>
<p><strong>Direct application control</strong>
The most innovative aspect is that AI directly controls various applications. This demonstrates AI’s evolution from a passive advisor or information provider into an active executor of real tasks.</p>
</li>
</ul>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="evas-approach-execution-centered-ai-in-physical-environments">EVA’s Approach: Execution-Centered AI in Physical Environments<a href="https://spectrabrain.ai/en/blog_tech/Innovation/openclaw#evas-approach-execution-centered-ai-in-physical-environments" class="hash-link" aria-label="Direct link to EVA’s Approach: Execution-Centered AI in Physical Environments" title="Direct link to EVA’s Approach: Execution-Centered AI in Physical Environments">​</a></h2>
<p><strong>EVA (Evolved Vision Agent)</strong> proves the same value of execution in the harsh realities of industrial environments.</p>
<ul>
<li>
<p><strong>Real-time site interpretation and visual reasoning</strong>
Through CCTV streams, EVA understands on-site context. It goes beyond object detection to answer higher-level questions such as:
“Why is that worker in danger?” or “Should a person be in that zone right now?”
This is a multimodal AI service built by optimizing Vision-Language Models (VLMs) for industrial environments.</p>
</li>
<li>
<p><strong>Direct physical response and control</strong>
EVA triggers physical actions in hazardous situations. Depending on severity, it can notify responsible personnel, activate on-site sirens, or even halt dangerous equipment processes.
AI does not stop at judgment — it functions as the final executor that prevents accidents.</p>
</li>
</ul>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="industry-wide-impact-a-major-ai-paradigm-shift">Industry-Wide Impact: A Major AI Paradigm Shift<a href="https://spectrabrain.ai/en/blog_tech/Innovation/openclaw#industry-wide-impact-a-major-ai-paradigm-shift" class="hash-link" aria-label="Direct link to Industry-Wide Impact: A Major AI Paradigm Shift" title="Direct link to Industry-Wide Impact: A Major AI Paradigm Shift">​</a></h2>
<p>The direction of AI development is entering a fundamentally different trajectory. In the past, progress meant building larger models with higher benchmark scores. Now, <strong>practical utility</strong> has taken center stage.</p>
<p>Beyond creating intelligent models, the key metric has become <strong>how completely AI can accomplish tasks within real user environments.</strong>
In other words, rather than model intelligence alone, the new compass for AI development is how deeply AI integrates into user workflows and delivers tangible value.</p>
<p>This paradigm shift is also redefining how we evaluate AI companies’ competitiveness. Moving beyond technological showmanship, market leadership will be determined by three core factors:</p>
<ul>
<li>
<p><strong>Execution reliability</strong>
Even highly intelligent systems must perform complex, multi-step tasks without interruption or error. Reliability will define trust.</p>
</li>
<li>
<p><strong>Integration capability</strong>
AI must seamlessly interoperate with existing software and systems rather than exist in isolation.</p>
</li>
<li>
<p><strong>User experience (UX)</strong>
Ultimately, success depends on how intuitively AI improves efficiency in real workflows and reduces user fatigue — not on technical flashiness.</p>
</li>
</ul>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="the-dawn-of-the-execution-centered-ai-era">The Dawn of the Execution-Centered AI Era<a href="https://spectrabrain.ai/en/blog_tech/Innovation/openclaw#the-dawn-of-the-execution-centered-ai-era" class="hash-link" aria-label="Direct link to The Dawn of the Execution-Centered AI Era" title="Direct link to The Dawn of the Execution-Centered AI Era">​</a></h2>
<p>The direction presented by OpenClaw represents more than technological advancement — it signals a paradigm shift across the AI industry. AI’s value will no longer be measured by <em>how intelligent it is</em>, but by <em>how useful it is in accomplishing real work</em>.</p>
<p>This transformation presents new opportunities and challenges for developers, enterprises, and users alike. Future AI competition will be decided not by benchmark scores, but by performance in real-world environments — marking a crucial turning point where AI truly becomes a tool that improves human life.</p>]]></content:encoded>
            <category>Tech</category>
            <category>Gen AI</category>
            <category>AI Agent</category>
        </item>
        <item>
            <title><![CDATA[Teaching VLMs to Multitask: Enhancing Situation Awareness through Scenario Decomposition]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Research/multi_scenario</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Research/multi_scenario</guid>
            <pubDate>Thu, 05 Feb 2026 18:00:00 GMT</pubDate>
            <description><![CDATA[At the core of EVA lies the ability to truly understand critical situations that occur simultaneously within a single scene—such as fires, people falling, or traffic accidents—without missing any of them.]]></description>
            <content:encoded><![CDATA[<p>At the core of EVA lies the ability to truly <strong>understand</strong> critical situations that occur simultaneously within a single scene—such as <em>fires</em>, <em>people falling</em>, or <em>traffic accidents</em>—without missing any of them.
However, no matter how capable a Vision-Language Model (VLM) is, asking it to reason about too many things at once leads to a sharp degradation in cognitive performance.[2,3]</p>
<p>In this post, inspired by the recent text-to-video retrieval research <strong>Q₂E (Query-to-Event Decomposition)[1]</strong>, we introduce <strong>Scenario Decomposition</strong>, a technique that enables VLMs to deeply understand <strong>complex, multi-scenario situations within a single frame</strong>.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="1--the-problem-cognitive-bottlenecks-in-vlms">1. 🚀 The Problem: Cognitive Bottlenecks in VLMs<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#1--the-problem-cognitive-bottlenecks-in-vlms" class="hash-link" aria-label="Direct link to 1. 🚀 The Problem: Cognitive Bottlenecks in VLMs" title="Direct link to 1. 🚀 The Problem: Cognitive Bottlenecks in VLMs">​</a></h2>
<p>EVA Agents are required to monitor complex urban roads and multi-use facilities 24/7.
In the early stages, we trusted the general-purpose capabilities of VLMs and expected them—like humans—to grasp all situations at once.</p>
<br>
<blockquote>
<p><strong>[Initial Approach: Unified Query]</strong>
"Look at this CCTV footage and check whether there is a fire, a fallen person, a traffic accident, or an emergency vehicle entering."</p>
</blockquote>
<br>
<p>In practice, however, the model revealed the limits of <strong>selective perception</strong>.
For example, it may clearly recognize a large bus in the center of the frame, while treating <strong>fire indicators</strong> or a <strong>fallen person</strong> behind it as mere background noise.</p>
<p>This occurs because the model’s attention mechanism is spread across multiple targets, resulting in shallow <strong>contextual understanding</strong> of critical risk signals.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="2--the-solution-extending-cognition-with-qe">2. 💡 The Solution: Extending Cognition with Q₂E<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#2--the-solution-extending-cognition-with-qe" class="hash-link" aria-label="Direct link to 2. 💡 The Solution: Extending Cognition with Q₂E" title="Direct link to 2. 💡 The Solution: Extending Cognition with Q₂E">​</a></h2>
<p>To address this issue, we adopted the core philosophy of Q₂E:
<strong>“Complex events become easier to understand when they are decomposed.”</strong></p>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="21-key-insight-from-qe-query-to-event">2.1. Key Insight from Q₂E (Query-to-Event)<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#21-key-insight-from-qe-query-to-event" class="hash-link" aria-label="Direct link to 2.1. Key Insight from Q₂E (Query-to-Event)" title="Direct link to 2.1. Key Insight from Q₂E (Query-to-Event)">​</a></h3>
<p>The Q₂E paper demonstrates that instead of searching with a single term like <em>“wildfire”</em>, decomposing it into an event flow—<strong>precursor → progression → outcome</strong>—allows the model to retrieve and understand the event far more effectively.</p>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="22-eva-agent-adoption-scenario-decomposition">2.2. EVA Agent Adoption: Scenario Decomposition<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#22-eva-agent-adoption-scenario-decomposition" class="hash-link" aria-label="Direct link to 2.2. EVA Agent Adoption: Scenario Decomposition" title="Direct link to 2.2. EVA Agent Adoption: Scenario Decomposition">​</a></h3>
<p>We extended this idea into <strong>Multi-Scenario Understanding</strong>.
Rather than asking a VLM to observe everything broadly at once, we assign it <strong>distinct perspectives</strong>, each focused on a specific scenario.</p>
<br>
<ul>
<li>
<p><strong>Before (Single Perspective):</strong>
“Notify me if there is a fire, a fallen person, or a traffic accident.”
→ (Ambiguous, attention dispersed)</p>
</li>
<li>
<p><strong>After (Decomposed Perspectives):</strong></p>
<ol>
<li><strong>[Fire Perspective]:</strong>
“Focus on pixel changes, colors, and smoke textures to identify signs of fire.”</li>
<li><strong>[Safety Monitoring Perspective]:</strong>
“Analyze human pose and interactions with surrounding objects to detect falls.”</li>
<li><strong>[Traffic Monitoring Perspective]:</strong>
“Examine vehicle collisions or abnormal stopping behavior to determine accidents.”</li>
</ol>
</li>
</ul>
<br>
<p>By decomposing the query this way, the VLM’s visual encoder activates <strong>strong, scenario-specific attention</strong>.
Even when viewing the same image, the model gains clarity on <strong>what it should perceive</strong>, enabling it to detect previously overlooked situations.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="3-️-system-design-a-specialized-multi-agent-architecture">3. 🏛️ System Design: A Specialized Multi-Agent Architecture<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#3-%EF%B8%8F-system-design-a-specialized-multi-agent-architecture" class="hash-link" aria-label="Direct link to 3. 🏛️ System Design: A Specialized Multi-Agent Architecture" title="Direct link to 3. 🏛️ System Design: A Specialized Multi-Agent Architecture">​</a></h2>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/figure1-84150d04fff1f199d44bfe403222d3da.png" width="90%"></div>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="overall-pipeline">Overall Pipeline<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#overall-pipeline" class="hash-link" aria-label="Direct link to Overall Pipeline" title="Direct link to Overall Pipeline">​</a></h4>
<p>The EVA Agent processes complex scenarios through the following five stages:</p>
<ol>
<li><strong>Classifier (Single or Multi):</strong> Determines the number of scenarios in the input</li>
<li><strong>Scenario Decomposition:</strong> Splits multi-scenario inputs into individual scenarios</li>
<li><strong>Scenario Enrichment:</strong> Injects rich detection conditions into each scenario [4]</li>
<li><strong>Scenario Clustering:</strong> Groups similar scenarios when the count exceeds four</li>
<li><strong>Parallel Inference:</strong> Executes VLM inference in parallel and aggregates results
<em>(Conceptual illustration: a single image processed through four separate prompt paths)</em></li>
</ol>
<br>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="stage-1-scenario-classification"><strong>Stage 1: Scenario Classification</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#stage-1-scenario-classification" class="hash-link" aria-label="Direct link to stage-1-scenario-classification" title="Direct link to stage-1-scenario-classification">​</a></h4>
<p>The LLM analyzes the user-provided scenario description to determine whether it represents a single or multiple scenarios.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[e.g.]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">"If a fire occurs, or someone falls, or a traffic accident is detected"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[LLM Prompt]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Analyze the following scenario description and determine whether it is</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">a single scenario or multiple scenarios:</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">"{user_input}"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[LLM Output]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">{</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  "type": "multi",</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  "count": 3,</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">  "scenarios": ["fire", "fall", "traffic accident"]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">}</span><br></span></code></pre></div></div>
<p><strong>Comparison with Q₂E:</strong></p>
<ul>
<li>Q₂E: Performs event decomposition from a single query like “2025 LA Fire”</li>
<li>EVA: Performs scenario decomposition from compound queries like “fire or fall or traffic accident”</li>
</ul>
<br>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="stage-2-scenario-decomposition"><strong>Stage 2: Scenario Decomposition</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#stage-2-scenario-decomposition" class="hash-link" aria-label="Direct link to stage-2-scenario-decomposition" title="Direct link to stage-2-scenario-decomposition">​</a></h4>
<p>Just as Q₂E decomposes a complex query into <em>Prequel / Current / Sequel</em>,
we decompose complex scenes into <strong>domain-specific directives</strong>.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[Q₂E Decomposition]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Query: "2025 LA Fire"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">→ Prequel: "What could happen before?"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">→ Current: "What happens during?"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">→ Sequel: "What could be the outcome?"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[EVA Decomposition]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Query: "fire or fall or traffic accident"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">→ Scenario 1: "Fire indicator detection"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">→ Scenario 2: "Fall accident detection"</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">→ Scenario 3: "Traffic accident detection"</span><br></span></code></pre></div></div>
<br>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="stage-3-scenario-enrichment"><strong>Stage 3: Scenario Enrichment</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#stage-3-scenario-enrichment" class="hash-link" aria-label="Direct link to stage-3-scenario-enrichment" title="Direct link to stage-3-scenario-enrichment">​</a></h4>
<p><strong>Comparison with Q₂E Refinement:</strong></p>
<table><thead><tr><th>Component</th><th>Q₂E</th><th>EVA</th></tr></thead><tbody><tr><td><strong>Input</strong></td><td>Decomposed Event<br>("Building on Fire")</td><td>Decomposed Scenario<br>("Fire indicator detection")</td></tr><tr><td><strong>Added Information</strong></td><td>Temporal (2025)<br>+ Spatial (LA)<br>+ Event (Fire)</td><td>Domain Knowledge<br>(visual features, detection criteria)</td></tr><tr><td><strong>Output</strong></td><td>"Building on Fire<br>during 2025 LA Fire"</td><td>"Fire detection via<br>refined detection criteria"</td></tr><tr><td><strong>Purpose</strong></td><td>Leverage LLM event knowledge</td><td>Focus VLM visual attention</td></tr></tbody></table>
<br>
<br>
<p><strong>Enrichment Example</strong>
(See also: <a href="https://spectrabrain.ai/en/blog_tech/Research/uid">Turning Simple User Requests into AI-Understandable Instructions</a>)</p>
<table><thead><tr><th>Aspect</th><th>Before (Simple)</th><th>After (Enriched)</th></tr></thead><tbody><tr><td><strong>Prompt</strong></td><td>"Detect fire."</td><td>"Detection: Fire is present with visible smoke.<br><br>Exception: Fire or smoke cannot be clearly identified due to camera angle."</td></tr><tr><td><strong>VLM Behavior</strong></td><td>Shallow scan of entire frame</td><td>Focused attention on color and texture features</td></tr></tbody></table>
<br>
<br>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="stage-4-scenario-clustering-when-n--4"><strong>Stage 4: Scenario Clustering (When N &gt; 4)</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#stage-4-scenario-clustering-when-n--4" class="hash-link" aria-label="Direct link to stage-4-scenario-clustering-when-n--4" title="Direct link to stage-4-scenario-clustering-when-n--4">​</a></h4>
<p>When the number of enriched prompts exceeds four, <strong>semantically similar scenarios are grouped</strong> to improve inference efficiency.</p>
<p><strong>Clustering Rule:</strong></p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">IF N &gt; 4:</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    Ask LLM to group scenarios by visual similarity</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    into a maximum of four groups</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">ELSE:</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    Process each scenario independently</span><br></span></code></pre></div></div>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[Example with many scenarios]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Input: ["fire", "smoke", "explosion", "fall", "collapsed person",</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        "traffic accident", "collision", "emergency vehicle"]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[LLM-based Semantic Grouping]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Group 1 (Fire-related): ["fire", "smoke", "explosion"]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Group 2 (Human safety): ["fall", "collapsed person"]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Group 3 (Traffic): ["traffic accident", "collision"]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Group 4 (Emergency response): ["emergency vehicle"]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">→ Reduce 8 scenarios to 4 groups</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">→ Optimize VLM latency and memory usage</span><br></span></code></pre></div></div>
<br>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="stage-5-parallel-vlm-inference"><strong>Stage 5: Parallel VLM Inference</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#stage-5-parallel-vlm-inference" class="hash-link" aria-label="Direct link to stage-5-parallel-vlm-inference" title="Direct link to stage-5-parallel-vlm-inference">​</a></h4>
<p>Each enriched prompt is sent to the VLM in parallel.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[e.g.]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">enriched_prompts = [Group1_prompt, Group2_prompt, Group3_prompt, Group4_prompt]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">PARALLEL_EXECUTE:</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">    FOR each prompt IN enriched_prompts:</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        result = VLM(frame=cctv_image, instruction=prompt)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">        results.append(result)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">WAIT_ALL_COMPLETE</span><br></span></code></pre></div></div>
<p><strong>Key Differences:</strong></p>
<ul>
<li>Q₂E analyzes a single video across multiple modalities (video, audio, text)</li>
<li>EVA analyzes a single frame across multiple scenarios (fire, fall, traffic)</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="4--case-study-unified-vs-decomposed-scenarios">4. 📊 Case Study: Unified vs. Decomposed Scenarios<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#4--case-study-unified-vs-decomposed-scenarios" class="hash-link" aria-label="Direct link to 4. 📊 Case Study: Unified vs. Decomposed Scenarios" title="Direct link to 4. 📊 Case Study: Unified vs. Decomposed Scenarios">​</a></h2>
<p>This is not about asking more questions—it is about <strong>how accurately the model understands the situation</strong>.</p>
<p>Using real EVA evaluation footage, we compared <strong>unified scenarios</strong> and <strong>decomposed multi-scenarios</strong> on the same complex video.</p>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="41-test-scenario">4.1. Test Scenario<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#41-test-scenario" class="hash-link" aria-label="Direct link to 4.1. Test Scenario" title="Direct link to 4.1. Test Scenario">​</a></h3>
<p>A video containing three critical incidents played sequentially was evaluated using both approaches.</p>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/figure2_eng-44a254b1403eff42622ee82eb2015df7.png" width="90%"></div>
<br>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="unified-scenario-prompt"><strong>Unified Scenario Prompt</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#unified-scenario-prompt" class="hash-link" aria-label="Direct link to unified-scenario-prompt" title="Direct link to unified-scenario-prompt">​</a></h4>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">"If at least one emergency vehicle exists,</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">or a fire occurs,</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">or at least one person has fallen"</span><br></span></code></pre></div></div>
<br>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="decomposed-multi-scenario-prompts"><strong>Decomposed Multi-Scenario Prompts</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#decomposed-multi-scenario-prompts" class="hash-link" aria-label="Direct link to decomposed-multi-scenario-prompts" title="Direct link to decomposed-multi-scenario-prompts">​</a></h4>
<p>Each scenario includes enriched detection criteria and exceptions.</p>
<table><thead><tr><th>Scenario</th><th>Detection Steps</th><th>Exceptions</th></tr></thead><tbody><tr><td><strong>Emergency Vehicle</strong></td><td>• At least one vehicle present<br>• At least one emergency vehicle identified</td><td>• No emergency vehicles<br>• All vehicles are taxis<br>• Police lights appear blue only<br>• Lights not mounted on roof<br>• Only headlights or brake lights visible</td></tr><tr><td><strong>Fall / Traffic Accident</strong></td><td>• At least one person present<br>• Person is fallen or traffic accident occurs</td><td>• No fallen person<br>• Only upper body visible<br>• Lower body not visible</td></tr><tr><td><strong>Fire / Smoke</strong></td><td>• Fire present<br>• Smoke detected</td><td>• No fire or smoke detected</td></tr></tbody></table>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="42-experimental-results">4.2. Experimental Results<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#42-experimental-results" class="hash-link" aria-label="Direct link to 4.2. Experimental Results" title="Direct link to 4.2. Experimental Results">​</a></h3>
<table><thead><tr><th style="text-align:center">Scenario</th><th style="text-align:center">Unified Prompt</th><th style="text-align:center">Decomposed Prompts</th></tr></thead><tbody><tr><td style="text-align:center">Fallen Person</td><td style="text-align:center">✅ Detected</td><td style="text-align:center">✅ Detected</td></tr><tr><td style="text-align:center">Emergency Vehicle</td><td style="text-align:center">✅ Detected</td><td style="text-align:center">✅ Detected</td></tr><tr><td style="text-align:center">Fire</td><td style="text-align:center">❌ Missed</td><td style="text-align:center">✅ Detected</td></tr><tr><td style="text-align:center">Overall Accuracy</td><td style="text-align:center"><strong>66.7% (2/3)</strong></td><td style="text-align:center"><strong>100% (3/3)</strong></td></tr></tbody></table>
<br>
<h3 class="anchor anchorWithStickyNavbar_LWe7" id="43-detailed-analysis">4.3. Detailed Analysis<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#43-detailed-analysis" class="hash-link" aria-label="Direct link to 4.3. Detailed Analysis" title="Direct link to 4.3. Detailed Analysis">​</a></h3>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="1-limitations-of-unified-scenarios-attention-interference"><strong>1) Limitations of Unified Scenarios (Attention Interference)</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#1-limitations-of-unified-scenarios-attention-interference" class="hash-link" aria-label="Direct link to 1-limitations-of-unified-scenarios-attention-interference" title="Direct link to 1-limitations-of-unified-scenarios-attention-interference">​</a></h4>
<p><strong>VLM Output</strong></p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Analysis:</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">✅ One fallen person detected (Scene 1)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">✅ One emergency vehicle (police car) detected (Scene 2)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">❌ No fire indicators detected</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Result: One critical incident missed</span><br></span></code></pre></div></div>
<p><strong>Issues:</strong></p>
<ul>
<li>Attention bias toward early scenes (people, vehicles)</li>
<li>Subtle fire indicators (smoke, flame textures) ignored in favor of salient objects</li>
</ul>
<br>
<h4 class="anchor anchorWithStickyNavbar_LWe7" id="2-decomposed-scenario-detection-result"><strong>2) Decomposed Scenario Detection Result</strong><a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#2-decomposed-scenario-detection-result" class="hash-link" aria-label="Direct link to 2-decomposed-scenario-detection-result" title="Direct link to 2-decomposed-scenario-detection-result">​</a></h4>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-background-color:hsl(230, 1%, 98%);--prism-color:hsl(230, 8%, 24%)"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="background-color:hsl(230, 1%, 98%);color:hsl(230, 8%, 24%)"><code class="codeBlockLines_e6Vv"><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">[Fire Detection Result]</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">⚠️ Fire detected</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Detected visual features:</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">✓ Orange/red flames (Scene 3, bottom-right)</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">✓ Vertical gray smoke patterns</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">✓ Irregular bright region boundaries</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain" style="display:inline-block"></span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Location: Bottom-right of frame</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Confidence: 0.91</span><br></span><span class="token-line" style="color:hsl(230, 8%, 24%)"><span class="token plain">Decision: Fire confirmed (immediate dispatch required)</span><br></span></code></pre></div></div>
<p><strong>Why it worked:</strong></p>
<ul>
<li><strong>Specialization:</strong> Fire Agent focuses exclusively on fire-related features</li>
<li><strong>Interference Reduction:</strong> Other scenarios do not introduce perceptual noise</li>
</ul>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="5-conclusion-the-depth-of-the-question-determines-the-depth-of-understanding">5. Conclusion: “The Depth of the Question Determines the Depth of Understanding”<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#5-conclusion-the-depth-of-the-question-determines-the-depth-of-understanding" class="hash-link" aria-label="Direct link to 5. Conclusion: “The Depth of the Question Determines the Depth of Understanding”" title="Direct link to 5. Conclusion: “The Depth of the Question Determines the Depth of Understanding”">​</a></h2>
<p>This work leads to a clear conclusion.
What limits VLM performance is not the number of parameters, but <strong>how we instruct the model to perceive the world</strong>.</p>
<p>Just as Q₂E improves retrieval accuracy by decomposing queries, we demonstrated that <strong>Scenario Decomposition</strong> enables VLMs to <strong>simultaneously and deeply understand multiple situations</strong> in complex CCTV monitoring environments.</p>
<p>Moving forward, we plan to extend this approach beyond static image analysis toward <strong>Temporal Event Decomposition</strong>, enabling models to reason over the temporal context of video streams.</p>
<p><em>(This post presents a creative adaptation of the Q₂E: Query-to-Event Decomposition methodology to solve Multi-Scenario Cognition challenges in vision-based tasks.)</em></p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="references">References<a href="https://spectrabrain.ai/en/blog_tech/Research/multi_scenario#references" class="hash-link" aria-label="Direct link to References" title="Direct link to References">​</a></h2>
<p>[1] Shubhashis Roy Dipta, Francis Ferraro.
"Q₂E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval."
arXiv:2506.10202v2, 2025.</p>
<p>[2] Standley, T., Zamir, A., Chen, D., Guibas, L., Malik, J., &amp; Savarese, S.
"Which Tasks Should Be Learned Together in Multi-task Learning?" ICML 2020.</p>
<p>[3] Pratt, S., Covert, I., Liu, R., &amp; Farhadi, A.
"What Does CLIP Know About a Red Circle? Visual Prompt Engineering for VLMs." ICCV 2023.</p>
<p>[4] EVA Tech Blog: Turning Simple User Requests into AI-Understandable Instructions</p>]]></content:encoded>
            <category>Tech</category>
            <category>Research</category>
            <category>EVA</category>
            <category>VLM</category>
            <category>AI Agent</category>
        </item>
        <item>
            <title><![CDATA[Physical AI Implemented with EVA]]></title>
            <link>https://spectrabrain.ai/en/blog_tech/Innovation/pat</link>
            <guid>https://spectrabrain.ai/en/blog_tech/Innovation/pat</guid>
            <pubDate>Thu, 22 Jan 2026 19:00:00 GMT</pubDate>
            <description><![CDATA[When Can AI Intervene in the Real World?]]></description>
            <content:encoded><![CDATA[<h2 class="anchor anchorWithStickyNavbar_LWe7" id="when-can-ai-intervene-in-the-real-world">When Can AI Intervene in the Real World?<a href="https://spectrabrain.ai/en/blog_tech/Innovation/pat#when-can-ai-intervene-in-the-real-world" class="hash-link" aria-label="Direct link to When Can AI Intervene in the Real World?" title="Direct link to When Can AI Intervene in the Real World?">​</a></h2>
<p>Accidents in industrial environments happen without warning.
Moments such as a worker collapsing, an arm getting caught in machinery,
or a fire breaking out usually occur within seconds.</p>
<p>Physical AI should not stop at recognizing these moments.
It must be capable of <strong>translating perception into physical action on site.</strong></p>
<p>In this post, we walk through a LEGO-based simulation to show
how EVA detects incidents
and how its decisions are connected to real equipment actions
as a single, continuous flow.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="simplifying-industrial-scenarios-with-lego">Simplifying Industrial Scenarios with LEGO<a href="https://spectrabrain.ai/en/blog_tech/Innovation/pat#simplifying-industrial-scenarios-with-lego" class="hash-link" aria-label="Direct link to Simplifying Industrial Scenarios with LEGO" title="Direct link to Simplifying Industrial Scenarios with LEGO">​</a></h2>
<p>Instead of replicating complex industrial environments in full detail,
we simplified accident scenarios using LEGO.</p>
<p>We designed independent scenarios for:</p>
<ul>
<li>a worker collapsing,</li>
<li>an arm being caught in equipment,</li>
<li>and a fire breaking out.</li>
</ul>
<div class="div_center"><p><strong>Arm caught in equipment – conveyor belt stops and warning light activates</strong></p></div>
<br>
<div class="div_center"><p><strong>Worker collapse – warning light and buzzer activated</strong></p></div>
<br>
<div class="div_center"><p><strong>Fire detected – conveyor belt stops and warning light activates</strong></p></div>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="eva-interpreting-situations-as-events">EVA: Interpreting Situations as Events<a href="https://spectrabrain.ai/en/blog_tech/Innovation/pat#eva-interpreting-situations-as-events" class="hash-link" aria-label="Direct link to EVA: Interpreting Situations as Events" title="Direct link to EVA: Interpreting Situations as Events">​</a></h2>
<p>What matters in this simulation is not simply
detecting a person or recognizing flames.</p>
<p>EVA interprets each situation through
predefined <strong>detection scenarios</strong>
and evaluates them as <strong>meaningful events</strong>.</p>
<p>Below is the interface where detection scenarios are configured in EVA.</p>
<div class="div_center"></div>
<p>Detection immediately becomes a <strong>trigger condition</strong>
for deciding the next action.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="physical-action-trigger-when-ai-decisions-move-reality">Physical Action Trigger: When AI Decisions Move Reality<a href="https://spectrabrain.ai/en/blog_tech/Innovation/pat#physical-action-trigger-when-ai-decisions-move-reality" class="hash-link" aria-label="Direct link to Physical Action Trigger: When AI Decisions Move Reality" title="Direct link to Physical Action Trigger: When AI Decisions Move Reality">​</a></h2>
<p>When an event occurs,
EVA’s role does not end at detection.</p>
<p>EVA transforms detected events into
<strong>Physical Action Triggers</strong>,
connecting them directly to on-site equipment and devices
so they can respond immediately.</p>
<p>The key point is that
each accident scenario is mapped to a predefined physical response.
Without waiting for human intervention,
AI decisions are translated directly into real-world actions.</p>
<p>Through this structure,
AI judgments do not remain as on-screen alerts or logs.
They become <strong>actions that actively change the state of the现场</strong>.</p>
<p>Physical Action Triggers represent the point where AI moves beyond
“what it sees”
to executing <strong>what must change in the real world</strong>.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="eva--n8n--equipment-control-workflow">EVA → n8n → Equipment Control Workflow<a href="https://spectrabrain.ai/en/blog_tech/Innovation/pat#eva--n8n--equipment-control-workflow" class="hash-link" aria-label="Direct link to EVA → n8n → Equipment Control Workflow" title="Direct link to EVA → n8n → Equipment Control Workflow">​</a></h2>
<p>These Physical Actions are not implemented with hardcoded logic.
They are built using a workflow-based approach.</p>
<p>Detection events generated by EVA are delivered to n8n via Webhooks.
Based on the severity and context of the event,
an Agent within n8n sends the appropriate control signals
to on-site equipment.</p>
<br>
<div class="div_center"><img src="https://spectrabrain.ai/en/assets/images/workflow-b3be821ecbd666b43257173b12ffd602.png" width="80%"></div>
<br>
<p>With this structure,
even if equipment changes or scenarios expand,
workflows can be reused and adapted flexibly.</p>
<br>
<hr>
<br>
<h2 class="anchor anchorWithStickyNavbar_LWe7" id="a-structure-that-makes-physical-ai-tangible">A Structure That Makes Physical AI Tangible<a href="https://spectrabrain.ai/en/blog_tech/Innovation/pat#a-structure-that-makes-physical-ai-tangible" class="hash-link" aria-label="Direct link to A Structure That Makes Physical AI Tangible" title="Direct link to A Structure That Makes Physical AI Tangible">​</a></h2>
<p>This LEGO simulation does not replicate
a real industrial site in full detail.</p>
<p>However, the structure—
where an incident occurs,
AI perceives it,
and a decision leads to a physical action—
is identical to real-world environments.</p>
<p>EVA does not leave AI as a result on a screen.
It enables AI to directly intervene in the physical world,
realizing the concept of <strong>Physical AI</strong>.</p>]]></content:encoded>
            <category>Tech</category>
            <category>EVA</category>
            <category>AI Agent</category>
            <category>Physical AI</category>
        </item>
    </channel>
</rss>