Reviewing binned Compensator Based Inference
My Review
The results described in this paper are useful and significant. One reservation I have is that this draft is very close to the previously published "Banerjee and Algeri (2026)". The submitted draft extends that method — derived for independent and identically distributed data — to binned histograms with Poisson counts. The conceptual contribution is largely inherited from the published work, and, as I explain below, the extension also weakens the main motivation for the method. So its not clear to me if this warrants a separate standalone publication. I have one substantive critique specific to the binned case, followed by minor comments.
The original paper shows that the effect of a misspecified background on the signal hypothesis test is captured by a single scalar — the compensator — equal to the component of the background misspecification aligned with the signal. When a background-only dataset is available the compensator is identifiable, and a test for the presence of signal can be constructed that controls the type I error rate for essentially any choice of the postulated background, however poor, subject only to mild support/integrability conditions.
The submitted draft shows that the same can be done with binned datasets.
My concern is that, in the binned case, the advantage of the compensator is unclear. Section 3 assumes a binned background-only histogram, which is what identifies the compensator. But a binned histogram is already a nonparametric description of the background: unlike the iid case, there is no arbitrary functional form to impose on fb, and hence none of the shape-misspecification bias that the compensator is designed to absorb. If binned fb and fs are available, η in Eq. (1) can be estimated directly by a binned maximum-likelihood fit, with the presence of signal assessed by the profile likelihood-ratio test. With a correctly specified binned background the model is correctly specified, the standard boundary asymptotics (½χ²₀ + ½χ²₁) apply, and there is no misspecification to compensate — the protective role of the compensator, which is the entire point of the iid construction, is unnecessary precisely in the regime where fb is directly available.
Of course the background-only histogram is a finite sample of fb rather than the truth, so a fair comparison is against a conventional binned fit that propagates the control-sample statistics in the usual way. Both approaches use the same finite control histogram; the question is which uses it more efficiently. The manuscript does not address this. I would ask the authors to (i) state explicitly what advantage the compensator provides over a standard binned template fit in the setting of Section 3, and (ii) support the claim with a direct comparison, especially in the relevant low-count-per-bin regime. With out this, it is not clear that the compensator buys anything once the data are binned.
This critique does not apply to Section 4, where no background-only sample is available and fb cannot be used directly. However, as discussed above, Section 4 closely parallels the corresponding treatment in the already-published iid paper.
Other than that I am happy with the submitted draft. A few minor comments below.
Logs
21 July 2026 Tuesday
- Finished.
- Submitted!
20 July 2026 Monday
- Figure 1: Should show the poisson errors on gray dots… thats the whole point of this extension. Would also be good to show binwidths with an x-error bar.
- Table 1: worth quoting the 68% CI on eta?
- p-values are quoted with precision of 10e-10 … seems very precise, how where these obatained ?
- would also be worth quoting the compensator values for the different postulated backgrounds and binnings?
eq1) worth defining E[N], E[ni] ? eq 4) worth define ||S||G ? Assume ||S||G2 is intergalX S2 dG(x)
16 July 2026 Thursday
[X]Setup page
TB: Signal Detection under Background Uncertainty
Use of AI in peer review To protect authors’ rights, reviewers must not upload a submitted manuscript or any part of it into an AI tool.
AI tools may be used in a supportive capacity only, for example to help improve the clarity or structure of review reports or to support background literature searches. Reviewers must ensure that all manuscript information remains confidential when using such tools. AI use must be disclosed in the review report, and reviewers remain fully responsible for the content of their review. See our Generative AI policy for more information.
Suggested disclosure statement:
During the preparation of this report, I used [NAME OF TOOL / SERVICE] in order to [REASON]. After using this tool/service, I reviewed and edited the content as needed and I take full responsibility for its content.
Previous paper
By studying the geometry of the problem, this article demonstrates that estimating the background distribution is somewhat unnecessary for inferring the signal intensity.