दोहा में Farah Rahman के सामने एक जाना-पहचाना मुद्दा था: एक ग्राहक बग रिपोर्ट, आधी रात की Slack थ्रेड, और 43 repositories में फैली 1.8 million lines वाली codebase। उसने एक agent से payment failure ट्रेस करने को कहा। पहली कोशिश ने इतने tokens खा लिए कि model ने संदर्भ खो दिया। फिर उसने आधा-सही जवाब बड़े आत्मविश्वास के साथ लौटा दिया।
यही वह तरह की गड़बड़ी है जिसे Semble ठीक करने की कोशिश कर रहा है। इसका pitch - code search for agents that uses 98% fewer tokens than grep - लगभग शरारती-सा लगता है। लेकिन शीर्षक के नीचे एक गंभीर बदलाव है: code search अब सिर्फ़ text को तेज़ी से ढूँढ़ने के बारे में नहीं है। यह machines को source का वही छोटा-सा हिस्सा देने के बारे में है जिसकी उन्हें ज़रूरत है। ऐसे रूप में, जिस पर वे budget या context जलाए बिना reasoning कर सकें।
ईमानदारी से कहें तो, यह मायने रखता है क्योंकि पुराना workflow इंसानों के terminals स्कैन करने के लिए बना था। Agents skim नहीं करते। वे ingest करते हैं। और जब वे ठीक से ingest नहीं करते, तो हर downstream action और अस्थिर हो जाती है: patch generation, root-cause analysis, test selection, doc fixes, यहाँ तक कि security review भी।
यह अभी क्यों महत्वपूर्ण है

लगभग एक साथ दो चीज़ें बदलीं। पहली, teams ने LLMs को सीधे engineering workflows में डालना शुरू किया, IDE assistants से लेकर autonomous triage bots तक। दूसरी, context की cost बहुत स्पष्ट हो गई। समझ में आता है? बड़े prompts सिर्फ़ महंगे नहीं होते; वे नाज़ुक भी होते हैं। अगर किसी agent को monorepo inspect करना हो, तो कुछ खराब retrieval steps एक पूरी run बर्बाद कर सकते हैं।
उसी समय, search expectations भी बदल गए हैं। Developers को अभी भी grep जैसी precision चाहिए, लेकिन agents को exact string matching से ज़्यादा चाहिए। उन्हें semantic narrowing, symbol-aware grouping, और इतना structure चाहिए कि वे irrelevant snippets से hallucinate न करें। Semble ठीक इसी gap में बैठता है।
मुख्य विचार: सिर्फ़ engineer के लिए नहीं, model के लिए search

Traditional grep एक ही काम में शानदार है: text matching। लेकिन यह बहुत ही शाब्दिक भी है। अगर आपका bug किसी renamed function, किसी generated file, या ऐसे symbol में व्यक्त है जो पाँच wrappers के पीछे छिपा है, तो grep अंधा हो सकता है—जब तक कि इंसान को सही needle पता न हो।
Agents के लिए code search अलग तरह से काम करना चाहिए। इसे एक विशाल repository को compact, high-signal evidence में बदलना चाहिए। इसका मतलब है relevance की संभावना के आधार पर ranking। आसपास की संरचना लौटाना, और repeated boilerplate को काट देना, जिसे मानव आँख नज़रअंदाज़ कर दे, लेकिन model खुशी-खुशी tokens पर खर्च कर दे।
गहरे स्तर पर, बात यह है कि tokens अब एक scarce engineering resource हैं। अमूर्त रूप से नहीं। शाब्दिक रूप से। Agent को दिया गया हर extra file chunk latency, cost, और answer quality बदल सकता है। इसलिए 98% fewer tokens than grep का दावा दिलचस्प है: इसलिए नहीं कि grep खराब है, बल्कि इसलिए कि grep को कभी LLM diet plan के रूप में डिज़ाइन ही नहीं किया गया था।
व्यावहारिक रूप से, agentic code search के लिए एक बेहतर system आम तौर पर कुछ चीज़ें अच्छी तरह करता है:
- सिर्फ़ lines नहीं, symbols को index करता है। Functions, classes, imports, और references को agents raw text blobs की तुलना में summarize करना आसान होता है।
- सिर्फ़ match count नहीं, context के आधार पर rank करता है। फेल हो रहे test के पास मिला एक high-value hit, generated docs में 27 occurrences से ज़्यादा महत्वपूर्ण हो सकता है।
- आक्रामक रूप से compress करता है। Repeated license headers, vendored code, और duplicated constants को prompt पर हावी नहीं होना चाहिए।
- Structured snippets लौटाता है। File path, symbol name, span, और dependency hints model को पूरे tree को फिर से पढ़े बिना reasoning में मदद करते हैं।
- Traceability बनाए रखता है। अच्छा agent search auditable होता है। आपको देखना चाहिए कि कोई snippet क्यों चुना गया।
- Tooling के साथ अच्छी तरह काम करता है। सबसे अच्छी search layer CI, IDEs, और terminal workflows में फिट होती है, किसी अलग ritual की माँग नहीं करती।
अगर model यह नहीं बता सकता कि उसने किसी file को क्यों चुना, तो शायद आप उस file को edit करने के लिए उस पर पर्याप्त भरोसा नहीं करते।
एक सूक्ष्म ergonomics लाभ भी है। इंसान खुद को orient करने के लिए search का उपयोग करते हैं। Agents कार्रवाई generate करने के लिए search का उपयोग करते हैं। ये एक ही task नहीं हैं। एक developer 12 noisy hits पर नज़र डालकर भी समझ सकता है कि कहाँ जाना है। इसके विपरीत, एक agent अगर retrieval layer sloppy हो, तो पूरे patch में गलत धारणा को आत्मविश्वास से फैला सकता है।
इसलिए Semble के token reduction वाले दावे को कुछ बड़े संकेतक की तरह पढ़ना चाहिए: बेहतर retrieval hygiene। कम noise का मतलब है hallucination की कम सतह। कम noise का मतलब यह भी है कि model के पास code में मौजूद वास्तविक invariants देखने के लिए अधिक जगह है। वहीं से उपयोगी बदलाव आते हैं।
व्यवहार में यह कैसा दिखता है

Kenji Silva, Krakow, Poland में एक data analyst, अपनी टीम को 14 services और 9 SQL models से जुड़ी pricing pipeline debug करने में मदद कर रहे थे। एक conventional search ने एक error string के लिए 311 matches निकाले। एक token-aware search workflow ने working set को 18 snippets तक सीमित किया और agent के prompt size को 91% घटा दिया। उन्होंने कहा कि पहली उपयोगी diagnosis 29 मिनट के बजाय 4 मिनट में आ गई।
Rina Popescu, Tallinn, Estonia में एक operations lead, 6 internal repos में incident playbooks inspect करने के लिए एक agent का उपयोग कर रही थीं। agent templated markdown और duplicated runbooks से बार-बार भ्रमित हो रहा था। जैसे ही उनकी टीम ने boilerplate को deduplicate करने वाली code-search layer पर स्विच किया, bot की false escalations एक हफ्ते में 17 से घटकर 3 रह गईं, और on-call team ने उसके आधे alerts को नज़रअंदाज़ करना बंद कर दिया।
दोहा में वापस, Farah Rahman ने 248 test failures वाले payment service पर एक pilot चलाया। उनकी टीम ने agent से failures को root cause के आधार पर cluster करने को कहा। बेहतर search layer के साथ, agent ने उनमें से 193 को 5 patterns में group किया और सटीक file paths सामने लाया जो बदले थे। इससे मानव निर्णय की ज़रूरत खत्म नहीं हुई। लेकिन बहुत सारी अंधी छानबीन ज़रूर कम हो गई।
बचने योग्य सामान्य गलतियाँ

-
Token savings को पूरी बात मान लेना। कम token use अच्छा है, लेकिन यही product नहीं है। अगर search कम context और खराब evidence लौटाती है, तो आपने बस error को compress किया है। लक्ष्य छोटे prompts नहीं, अधिक भरोसेमंद decisions हैं।
-
Semantic tasks के लिए exact-match सोच का उपयोग करना। grep तब बहुत अच्छा है जब आपको string, symbol, या import path पता हो। Agents अक्सर नहीं जानते। अगर आपकी retrieval strategy बस grep को पतले wrapper के साथ दोहराती है, तो आप renamed code, implicit dependencies, और generated artefacts चूक जाएँगे।
-
Repository structure को अनदेखा करना। Test helper में function definition और production code में उसी नाम का function बराबर नहीं होते। Agents को ownership, module boundaries, और call direction के संकेत चाहिए। इनके बिना, वे गलत layer को भी पूरे आत्मविश्वास के साथ patch कर सकते हैं।
-
बहुत ज़्यादा boilerplate देना। License blocks, vendor bundles, minified JS, और repeated config fragments context को तेज़ी से खा सकते हैं। Teams अक्सर इन्हें नज़रअंदाज़ कर देते हैं क्योंकि इंसान मानसिक रूप से clutter के पार देख लेते हैं। Models skip नहीं करते; वे absorb करते हैं।
-
Evaluation छोड़ देना। Demo में clever दिखने वाली search layer long-tail naming, polyglot repos, या messy tests वाले real workloads पर फेल हो सकती है। Retrieval quality को concrete tasks से मापें: bug localisation, file ranking, और answer correctness, सिर्फ़ query latency से नहीं।
-
Assuming the agent will self-correct. अक्सर ऐसा नहीं होगा। अगर retrieval step गलत file cluster की ओर इशारा करता है, तो model उसके ऊपर एक साफ़-सुथरी लेकिन झूठी कहानी बना सकता है। अच्छी search पहली guardrail है, optional enhancement नहीं।
एक व्यावहारिक checklist

-
एक दर्दनाक workflow से शुरू करें। Bug triage, test failure analysis, या dependency tracing जैसे किसी task को चुनें। तब तक पूरी ocean को उबालने की कोशिश न करें जब तक आपको न पता हो कि कौन-सा search pain सबसे महत्वपूर्ण है।
-
पहले और बाद में prompt size मापें। Input tokens, output quality, और first usable answer तक का समय ट्रैक करें। अगर baseline का नाम नहीं ले सकते, तो सुधार साबित नहीं कर सकते।
-
हर snippet क्यों लौटाया गया, यह log करें। File path, symbol, और relevance signal रखें। जब model अजीब छलांग लगाता है, तो debuggability मायने रखती है।
-
Dead weight जल्दी हटाएँ। Vendored directories, build artefacts, और duplicated generated files को बाहर करें। यह सस्ता है और अक्सर तुरंत लाभ देता है।
-
Raw text dumps की बजाय structured retrieval को प्राथमिकता दें। Agent को एक compact package दें: symbol name, आसपास की lines, और dependency hints। यह आमतौर पर पूरे files paste करने से बेहतर होता है।
-
Ugly queries के साथ टेस्ट करें। Renamed functions, partial error messages, और “checkout thing breaks after retry” जैसी vague descriptions आज़माएँ। असली agents mess देखते हैं, perfect keywords नहीं।
-
grep से ईमानदारी से तुलना करें। बहुत से human tasks में grep अभी भी जीतता है। देखें कि नई search agent की कहाँ मदद करती है और पुराना tool कहाँ तेज़ और सरल रहता है।
-
False confidence पर नज़र रखें। अगर agent अधिक fluent लेकिन कम accurate हो जाता है, तो retrieval को कसें और source spans तक citations अनिवार्य करें।
यह कब नहीं करना चाहिए
हर टीम को agent-first search layer की ज़रूरत नहीं होती। अगर आपका repo छोटा है, आपका stack साफ़-सुथरा है, और आपके tasks ज़्यादातर मानव-चालित हैं, तो grep plus एक अच्छा IDE search पर्याप्त हो सकता है। जब असली bottleneck file ढूँढ़ना नहीं, बल्कि product को समझना हो, तब fancy retrieval दिखावा बन सकती है।
यदि आपके पास outputs का मूल्यांकन करने के लिए पर्याप्त operational discipline नहीं है, तो यह भी गलत कदम है। Token-efficient search engine experiments को सस्ता बना सकता है, जो अच्छा है, लेकिन यह bad agent behaviour को भी सस्ता बना सकता है। अगर कोई retrieval quality की समीक्षा नहीं कर रहा, तो आप बस confusion को कम लागत पर automate कर रहे होंगे।
और जानने के लिए
- https://owasp.org/www-project-top-ten/ - security की मूल बातें, जो तब भी मायने रखती हैं जब agents code से छेड़छाड़ करते हैं
- https://developer.mozilla.org/en-US/docs/Web/JavaScript/Guide/Regular_expressions - यह याद दिलाने के लिए कि grep-style searching वास्तव में अंदर से क्या कर रही है
- https://en.wikipedia.org/wiki/Information_retrieval - ranking, relevance, और search evaluation के पीछे का व्यापक क्षेत्र
जाहिर है, Semble का दिलचस्प हिस्सा marketing number नहीं है, भले ही 98% fewer tokens एक अच्छा headline हो। दिलचस्प संकेत यह है कि code search को पहले machine readers, फिर human readers के लिए फिर से बनाया जा रहा है, और यह बदलाव teams के debug, patch, और software automate करने के तरीके को बदल देगा। सवाल अब यह नहीं है कि agents code search कर सकते हैं या नहीं - सवाल यह है कि आपकी search layer उन्हें साफ़-साफ़ सोचने में मदद कर रही है या बस उन्हें और noise दे रही है।