How blocking an API in robots.txt broke Google’s view of our React app
By NeuralNews · Incident observations: 22–29 September 2026. Prepared with AI assistance from our implementation records and owner-provided Search Console screenshots.
The symptom: available to Google, but an empty feed
On 22 September, Search Console classified our AI news page as a soft 404. A live URL test said the URL was available to Google, yet its rendered screenshot showed “Failed to load news”, zero briefing picks, and an unavailable coverage status. The page was reachable, but its main content was missing. That distinction mattered more than the green availability check.
Our rendering path
NeuralNews uses Express and React. Express serves page metadata and a feed shell containing article links. At startup, React mounts with createRoot and replaces that shell; the client then fetches stories from /api/news. The server shell did not guarantee that the final rendered page would retain those stories. If client data loading failed, the crawler could end up with an error screen instead.
The rule that blocked a rendering dependency
Our production robots.txt disallowed the entire API prefix. We had treated JSON endpoints as something search engines should not visit. But a crawler rendering our React page needed those public read endpoints. This was a concrete configuration defect, supported by the failed rendered view. We did not capture Google’s individual API response codes, so we cannot attribute every historical indexing exclusion to this one rule.
User-agent: *
Allow: /
Disallow: /api/The fix: allow public reads, keep expensive routes excluded
We replaced the broad API block with exclusions for push, content extraction, and summarization routes. This made the news feed, individual articles, source list, metadata, and briefing requests crawlable. Existing rate limits remained in place. Robots directives guide cooperating crawlers; they are not authentication or access control.
User-agent: *
Allow: /
Disallow: /api/push/
Disallow: /api/news/content
Disallow: /api/news/summarize
Sitemap: https://www.neuralnews.in/sitemap.xmlCrawl permission and indexing permission are different
Our existing middleware sends X-Robots-Tag: noindex on API responses. Allowing those requests lets the renderer obtain data without inviting the JSON URLs into search results. The HTML pages retain their own indexing directives. A crawler must be able to fetch a response to discover its noindex instruction; blocking the URL in robots.txt is not a substitute.
app.use('/api', (_req, res, next) => {
res.set('X-Robots-Tag', 'noindex');
next();
});A regression test with a negative control
The regression test reads the served robots.txt and checks the public resources the client needs. It asserts that these requests are unblocked, return successfully, and retain the noindex header. We temporarily restored Disallow: /api/ and confirmed the test failed for the AI feed request, then restored the fix and confirmed it passed. Our test checks literal Disallow prefixes for this simple configuration; it is not a general robots.txt parser. Multiple user-agent groups, wildcard rules, or Allow overrides would require a standards-aware checker.
What changed in Search Console
The 22 September live test at the displayed time 22:57 showed the failed feed. Later tests at 23:07 for /ai-news and 23:09 for /tech-news displayed actual story cards and five briefing picks. A sample article also rendered at 23:10. On 29 September, the stored URL Inspection results for the homepage and /ai-news said “URL is on Google” and “Page is indexed”. These observations establish successful rendering for the tested pages and subsequent indexing of those two URLs. They do not establish that every article is indexed or that this change caused a traffic increase.
How to check the same failure on your site
First, inspect a representative URL in Search Console and run a live test. Open View tested page and inspect both the screenshot and rendered HTML for real headlines and links. Next, compare the initial HTML with the rendered page and identify every resource needed to load its main content. Check robots rules for those exact paths, as well as response status, authentication requirements, and rate limits. A successful curl request alone does not prove Google is allowed to fetch the resource: curl does not enforce robots.txt.
Verify the repair, then monitor the stored result
After changing robots.txt, account for crawler caching and repeat the live test. Confirm the main content is present before requesting indexing for important repaired pages. Record the stored URL Inspection crawl date, indexing decision, and selected canonical separately from the live test. Recheck after Google recrawls; repeated indexing requests are not a substitute for fixing a remaining rendering or content problem.
What this repair does not solve
Making content render removes a technical obstacle. It does not make short syndicated excerpts unique or guarantee rankings. Our next content work is to publish useful original explanations, such as this incident report, with attributable evidence. For another React application, preserving useful server content through hydration or a failed client refresh may also reduce dependence on a successful follow-up API request. That is an architectural option, not a claim that we implemented it in this repair.
Evidence and corrections
This report is based on our deployed robots configuration, middleware, regression-test checks, and the dated Search Console observations described above. The screenshots are not embedded here. We have not measured a causal traffic uplift or field Core Web Vitals improvement from this change. Send corrections or questions using the contact address below.