1. Classify the work
Each article identifies itself as a decision framework, implementation guide, research method, buyer guide or original empirical research. A framework synthesizes existing evidence and engineering practice; it is not a benchmark result. We label original testing only when we have run and documented it.
2. Start with primary material
We prefer papers, standards, official technical documentation, repositories, model cards and regulatory text. Vendor announcements can establish what a vendor claims, but not that the claim transfers to another workload. Material sources appear beside the article, with publisher and direct link.
3. Predeclare evaluation choices
For empirical work, the intended publication standard is to freeze tasks, candidates, configurations, metrics, stop conditions and scoring rules before inspecting final results. We record model and tool versions, dates, retries, intervention and exclusions. We do not currently publish a proprietary cross-vendor benchmark.
4. Measure systems, not names
Agent and model performance depends on prompts, tools, retrieval, permissions, infrastructure and reviewers. Analysis therefore describes the configured system and the operating context rather than attributing every outcome to a model alone.
5. Preserve uncertainty
Sample counts, missing evidence, transfer risk and material limitations belong in the result. Small samples are not converted into precise rankings. A result may support a prototype, a constrained deployment or no decision yet.
6. Corrections and updates
Articles show publication and update dates. Material corrections should explain what changed and why. Send reproducibility questions or correction requests to mann@neuralarc.in.