The Search Brief · News analysis
Bing’s AI Performance Report Gives Publishers a Citation Baseline—With Important Limits
Bing’s AI Performance report gives publishers a citation baseline. Learn how to interpret references, map pages to purpose and separate visibility from outcomes.
Development covered: February 10, 2026

Publishers have spent much of the AI search debate trying to infer visibility from scattered screenshots and referral traffic. Bing’s AI Performance report provides a more direct observation point: a record of citations within supported experiences. The opportunity is useful, provided teams do not convert citation counts into a ranking or revenue metric that the report does not claim to measure.
Microsoft’s February 10 announcement introduced the public preview in Bing Webmaster Tools for Microsoft Copilot, AI summaries in Bing and selected partner integrations. Its metrics include total citations, cited pages and sampled grounding queries. Microsoft distinguishes citation counts from ranking or placement, and grounding queries from the complete prompts users submit. Average cited pages is a daily average of unique pages, not a count of unique people. Those definitions should accompany any report derived from the interface.
The practical benefit is a baseline. A publisher can begin asking which resources are being referenced and whether that coverage matches its editorial priorities. It still needs other evidence to understand the accuracy of an answer, the value of a visit or the commercial outcome of the exposure.
A citation answers one question
A citation indicates that a source was referenced within the measured experience. That is useful information about visibility. It does not by itself tell the publisher whether the source was prominent, whether the user noticed it or whether the surrounding answer represented it correctly.
Those are separate questions that require separate observations. A highly visible citation might send few visits because the answer satisfies the immediate need. A less frequent citation on a complex comparison topic might lead to a valuable research session. The count alone cannot distinguish those situations.
This is why the first dashboard should remain modest. Show what is directly reported, preserve the definitions and identify the next question. A baseline that is honestly limited is more useful than a elaborate score that mixes unlike measurements.
For a directory, the first question might be whether important category and methodology pages appear at all. That observation can inform further investigation without implying that every citation is a recommendation of the directory’s commercial choices.
Map cited pages to their editorial purpose
A list of URLs becomes more useful when grouped by what each page is intended to do. A definition article, a category comparison, a service profile and a methodology page have different roles. Their citation patterns should not be evaluated against an identical expectation.
Create a simple inventory with the page’s topic, audience, purpose and maintenance owner. Add the observed citation data without changing the underlying definition. This makes it possible to see whether coverage is concentrated in a narrow group or distributed across important resources.
The exercise may reveal a mismatch. A site could be referenced for broad definitions while its detailed comparisons receive little attention. That does not automatically mean the comparisons are defective, but it identifies a question worth investigating. Are they accessible, clear, current and distinct from other resources?
Avoid treating every uncited page as a failure. Some pages serve existing customers, navigation or specialized needs with limited observed demand. The goal is to understand the portfolio, not force every URL to generate the same kind of exposure.
Grounding queries are research clues
A retrieval phrase can reveal how a system searched for supporting information, but it should not be treated as a verbatim transcript of a person’s entire request. That distinction matters when a marketer builds content briefs or claims to know exactly what customers asked.
Use the phrases to identify themes and possible information needs. Compare them with the page’s intended purpose and with other available research. If a phrase is ambiguous, retain that ambiguity rather than expanding it into a confident story about the user’s motivation.
A useful editorial question is whether the page supplies a clear, attributable answer to the apparent subproblem. For example, a resource about agency selection might explain how to compare scope or verify a case study. The page should not be rewritten to repeat every observed phrase mechanically.
The research value comes from recognizing patterns, not from treating a sample as a complete demand census. Keep the date and reporting scope with the export so later comparisons remain interpretable.
Build an accuracy review alongside the visibility review
If a publisher can observe an answer containing a citation, it should inspect whether the answer accurately represents the source. Check the claim, the destination page and any important qualification that may have been omitted. A citation to an outdated or ambiguous passage can be a reason to improve the page.
Record the observation conditions and avoid claiming that one answer is universal. Generative experiences can vary. The purpose of the review is to identify concrete problems, not to produce an absolute statement about every user’s experience.
For a directory, look carefully at the distinction between editorial placement, verified facts and user reviews. If those concepts are unclear on the page, an external summary may reproduce the ambiguity. Improving the visible explanation is a useful action even when the publisher cannot control every downstream answer.
| Review item | Question |
|---|---|
| Referenced claim | Does the page actually support it? |
| Qualification | Is an important limit easy to find? |
| Currency | Does the source still describe the current situation? |
| Attribution | Is a vendor claim distinguished from independent evidence? |
| Destination | Does the citation lead to the intended resource? |
Keep referrals and outcomes in a separate layer
Referral analytics can help describe visits that arrive from identifiable sources, but it does not necessarily capture every exposure or every later action. A person may return directly, use another device or take no action. Those limitations should remain visible in the reporting.
When visits are observed, examine what happens next. Do readers reach the relevant comparison, use a useful filter or follow a meaningful contact path? The definition of a valuable action should come from the business purpose of the page, not from whichever event is easiest to count.
Do not assign a monetary value to every citation by multiplying it by an arbitrary advertising rate. That can produce a convenient total without demonstrating actual business value. If an estimated value is used for planning, label it as a model and show its assumptions.
A clear report can place citation observations, identifiable visits and completed actions in adjacent sections. Keeping them separate allows readers to understand the relationship without implying a level of attribution that the data does not support.
A baseline should survive changes in page structure
Websites evolve. Pages are consolidated, slugs change and categories are reorganized. If reporting relies only on the current URL, historical comparisons can become confusing. Keep a record of redirects and consolidations so the team can interpret changes in cited-page counts.
For example, combining three overlapping articles into one stronger resource may reduce the number of distinct URLs while improving the usefulness of the destination. A decline in a page-count metric would not necessarily be a negative result. The editorial context matters.
Conversely, splitting one resource into many thin pages could increase the number of URLs without creating more value. A dashboard that rewards breadth alone can encourage that mistake. Evaluate whether the pages have distinct purposes and adequate substance.
The maintenance record should connect old and new URLs and note when the change occurred. This is ordinary publishing discipline, but it becomes especially valuable when a new visibility metric makes URL counts appear strategically important.
Test improvements with a specific mechanism
A reasonable experiment begins with an observed weakness. Perhaps a page buries a useful comparison, lacks a date for changing information or presents an important definition ambiguously. The proposed change should address that weakness directly.
Document the page, the change, the intended benefit and the observation period. Keep a comparable group unchanged where practical, while recognizing that a small editorial test rarely eliminates every confounding factor. The purpose is to improve the quality of inference, not to claim laboratory certainty.
If visibility improves, report the timing and limitations. If it does not, inspect whether the page became more useful anyway and whether the original hypothesis was sound. A failed visibility hypothesis can still reveal something about the workflow, but it should not be relabeled as a success without explanation.
Avoid making many unrelated changes and then attributing the outcome to one fashionable technique. Clearer structure, stronger evidence and repaired access can all matter to the page, but their individual effects may not be separable in a bundled release.
Do not let the report dictate the entire content strategy
A new dashboard naturally attracts attention. The risk is that the organization begins creating only the kinds of pages that appear easily in the report. That can narrow the editorial program around a measurement surface rather than the needs of the audience.
Some valuable resources support later stages of a decision or serve a small specialist audience. They may not generate large citation counts. Others can attract broad references without helping the business’s intended readers. Neither outcome should be judged solely by volume.
Use the report as one source of feedback within a wider strategy. Customer questions, sales conversations, support issues and direct readership can reveal needs that a citation dataset does not. A healthy program can explain why a page matters even before it accumulates a large number.
This is especially important for news analysis. An article should help readers understand a development and its implications. It should not merely be formatted to maximize the chance that a system extracts one sentence while the rest adds little value.
What a useful monthly review looks like
Begin with changes in observed coverage and the definitions needed to interpret them. Identify which page groups contributed to the change, and note any site releases or reporting changes that affect comparability. Then examine a small set of meaningful examples rather than presenting only an aggregate total.
For each example, connect the observation to a decision. A stale cited page may need updating. An important uncited resource may deserve an accessibility and content review. A well-represented topic may indicate that the existing explanation is useful and should be maintained.
Finish with a short list of actions, owners and review dates. The report should reduce uncertainty or change the work. If every month ends with the same generic recommendation to publish more, the team may not be using the new data effectively.
Keep a record of unresolved questions. Some will require more time, a different data source or a more carefully designed test. Stating that need is more honest than filling the gap with an invented explanation.
A direct observation point is progress
Bing’s report gives publishers a clearer place to begin studying AI citations. Its value is strongest when the organization respects the boundaries between a reference, a visit and an outcome. Those boundaries make the analysis more precise and the resulting work easier to evaluate.
The next step is practical: establish a baseline, map pages to purpose, inspect important examples and improve the underlying resources where a real weakness is visible. That creates a measurement program capable of learning, rather than another dashboard used mainly to support confident claims.
Make the first review small enough to complete
An initial citation review does not need to cover every URL at equal depth. Select a few important page groups and inspect the observations most likely to affect a decision. A manageable review can produce concrete repairs while the team learns how the report behaves.
For each selected page, keep the source observation, the editorial assessment and the proposed action together. If the page is current and clear, the right action may be continued observation. If an important qualification is missing, repair it before speculating about more elaborate optimization tactics.
Set a review date and identify what new evidence would matter. A longer observation period may resolve a low-volume pattern. A live answer sample may reveal a representation problem. A technical check may explain why a destination is unavailable. The next step should answer the unresolved question.
This sequence prevents a new reporting feature from creating an unlimited workload. The team learns from a defined set of cases, documents the result and expands the process where it proves useful. A completed, well-explained review is more valuable than a large export that never changes the website.
Source and analysis note: Report definitions are attributed to Microsoft’s announcement. The review process, examples and measurement recommendations are SEOS.co analysis. No citation increase or causal optimization result is claimed for SEOS.co in this article.