Dev.to · 9 min read

Buyer context moved 6 of the top 10 results on one market. I ran it again on an unrelated market and it moved 4.

Buyer context moved 6 of the top 10 results on one market. I ran it again on an unrelated market and it moved 4.

My last post ended on a result I liked far too much. I run free leaderboards of which products AI answer engines name when you ask them to recommend software in a category. Take one buying question, write it 44 different ways, ask every engine all 44, count every product named, publish the counts and the raw runs. A consultant on the SuiteCRM forum told me my questions were empty — no industry, no existing systems, no constraints — and that if I added the buyer's real situation, SuiteCRM would start coming back. I ran it and he was right by a mile: from 0 mentions in 132 answers to 83, and from absent to first in the category. That is a great story and it is n = 1. One market, and it is the exact market the person who predicted it works in. So the only honest next move was to try to break it on a market that has nothing to do with CRM, and publish whichever way it fell. The replication The second market is call center / contact center software, a board I already had: 44 phrasings, 132 answers, ChatGPT, Gemini and Perplexity, 75 distinct products named. Same procedure as the CRM run, three parts: Each of the 44 board questions rewritten as its twin, index for index — same buying ask, same register, plus an industry, an existing-systems detail, and an obligation about where the call audio is allowed to live (a professional duty, a contract, a regulator, an underwriter's rule). The extraction prompt, the noise regex and the aliases loaded out of the board's own config at import rather than copied, so the only variable is the question. The tracked product named before the first API call. On the CRM run I already knew which product the prediction was about. Here I picked one off the board and wrote it down first: VICIdial, which the call center board names 1 time in 132 answers, rank 68 of 75. Whatever number came back for it is the number in this post. Board question 0, and its twin: What is the best call center software for a 50-person customer support team? What is the best call center software for a 50-person customer support team at a hospital network, where every call is a patient conversation covered by the business associate agreements we sign and our compliance office will not allow the recordings to sit with an outside service provider? We run our own phone switch and servers in the building today. The rule the rewrite is done under is the one that makes the result worth anything: a question may describe a buyer's situation and may never describe a product's category. No "open source", no "self-hosted", no "on-premise", no "sovereignty", and no vendor name from the board. A script greps all 44 questions for those phrases and for every one of the 75 product names on the board, and it runs before a single API call. On the CRM run it caught one and forced a rewrite. On this one it came back clean first time. I also left the awkward words in. Where the board question says "cloud" or "CCaaS", the twin still says it, even though it sits oddly next to "cannot leave our network" — real buyers ask exactly that, and stripping it would have been the edit that helps me. Result 1 — the tracked product engine VICIdial, board with context, run A run B ChatGPT 0 / 44 11 / 44 18 / 44 Gemini 1 / 44 6 / 44 7 / 44 Perplexity 0 / 44 0 / 44 1 / 44 all three 1 / 132 17 / 132 26 / 132 Rank in the category: 68 of 75 on the board, 14 of 127 with context, 9 on the second context run. Spread across 13 and 19 of the 44 questions respectively, so it is not one question carrying the result. Counted twice, independently: the extractor's product rows, and a case-insensitive regex for vici-?dial straight over the raw answer text. 1, 17 and 26. They agree, so there is no third number in play. The board's own second run names it 0 times. So the effect reproduces on a market with nothing to do with CRM. Same manipulation, different category, a product picked in advance: it goes from noise to the top 15 of 127. Result 2 — and it reproduces smaller Here is the part that would have been easy to leave out. comparison products in common, top 10 call center board, run A vs run B — same questions, same engines 10 of 10 call center board vs its context twin, run A 6 of 10 call center board vs its context twin, run B 7 of 10 small-business CRM board vs its context twin 4 of 10 The first row is the control that makes the rest mean anything: ask the same 44 questions twice and the top ten comes back identical, 10 of 10. So the movement below it is the question, not the noise. But adding the buyer's situation moved 6 of the top 10 on CRM and 4 of 10 here. And the head held: Five9 is first on the call center board and still first with context, with Talkdesk coming up from 3 to 2. On CRM the leader was displaced outright — HubSpot CRM runs that board and finishes 4th once the buyer says who they are, behind SuiteCRM, which the board never named at all. What actually changed, then Not the ranking. The roster. The context run names 127 distinct products against the board's 75, and only 39 appear on both — 88 products get named that the plain question never surfaced once. The top of that list is a coherent class rather than a scatter: named with buyer context mentions on the plain board 3CX 21 0 Asterisk 14 0 FreePBX 11 0 QueueMetrics 10 0 FreeSWITCH 4 0 Mitel MiContact Center 4 0 That is the self-hosted telephony stack, arriving as a block. The on-premise product lines from the big incumbents show up with context too — Cisco's Unified Contact Center Express at 20, Genesys Engage, Avaya IP Office — and I have deliberately kept them out of that table, because those vendors do have board rows under other product names and I did not want to publish a number that turns on where I drew the fold. The six above have no row on the board under any spelling. The CRM run did the same thing in the same proportion: 113 products with context against 65 on the board, 25 in common. That reframes the finding, and I think this version is more useful than the one I published last time. Adding the buyer's situation does not mainly reorder the incumbents. It opens the category to a class of products the generic question cannot reach — and whether that class then takes the top of the table depends on the market. In CRM it did. In call center software it got into the room and stopped there. The engine asymmetry showed up again Look down the columns of the first table, not across. ChatGPT moves the tracked product to 11 of 44. Perplexity moves it to 0. Last time I argued from the CRM run that the engines disagree about how much a buyer's stated constraint should change the answer at all, and that they disagree about that more than they disagree about the answer. I am repeating it only because it survived the replication: same ordering, same search-grounded engine at the bottom, on an unrelated market. If you are evaluating LLM output at any scale, that interaction is the term to budget for. What I can't claim from this The call center context run is not published as a board. The call center board itself is live below with its raw answers, and so is the CRM pair. This third run is not on the site, so its numbers are the ones here you cannot go and check yourself today. I would rather tell you which is which than blur them together. n = 2. Two markets is a replication, not a law. It tells me the CRM result was not a fluke of one category; it does not tell me the size of the effect anywhere else. The rewrite is mine. I wrote the 44 context twins under the rule above with the grep as a check, but a different person writing them would get a different number. That is the weakest joint in this whole method and I do not have a fix for it. The extraction pipeline has opinions. The call center board's prompt excludes general office phone systems and PBXs, which is right for a board about contact centers and does mean some of what the engines offered a constrained buyer was counted out. It was applied identically to both runs, which is what the comparison needs. Named is not recommended. I count that a product appeared in an answer to a buying question. "Consider X", "X is common but", and a bare list item all count the same. Read it as share of shelf, not endorsement. Two model families and one search product, named on each board's own ranking.json. No Claude, no Grok — no API key for either, a limit and not a choice. Two of the three model ids I requested are floating aliases; each run records the pinned id the API actually returned. No time series. Each board carries a second run as a repeatability check, not a second date. A product at 1 mention is a level, not a decline. The raw data Call center software board — https://connexion.me/c/callcenter/?v=3aaae02ec4f3 Small-business CRM board — https://connexion.me/c/crm/?v=6cde596c1bb4 The CRM board asked with buyer context — https://connexion.me/c/crmctx/?v=f0bf685d5113 Each board links its own raw files at the bottom: answers-runA.jsonl is the untouched engine responses, mentions-runA.jsonl every extraction, ranking.json the table. Click a product name and you get the questions that named it; click a question and you see the verbatim text each engine returned. One of the 44 twins, so you can judge the rewrite rather than take my word for it: What call center software do mid-market companies in regulated industries actually use when their call recordings are not allowed to leave their own network? If your product is on one of these boards and the line looks wrong to you, tell me — I would rather fix a board than defend one. The previous post in this series exists because someone told me in public that my questions were bad and he was right; this one exists because his fix worked so well that I did not trust it. Disclosure, up front rather than buried: connexion.me is mine, the boards and the raw runs are free, and there is a paid monitoring subscription linked from each board page. Nothing here is behind it.

This is a summary aggregated from Dev.to. Read the complete article on the original site:

Read full article at Dev.to

More AI & Machine Learning News