Nonresonant triple Higgs boson production 6b
Group notes
Congratulations on the interesting paper. Given this is the first result in this channel I would rather see this coming out as a PRD-sytle paper than PRL, from which its hard to convey to reader what you really did.
L15 "Although HH production is primarily sensitive to κ3, it also depends on κ4 through electroweak (EW) corrections; however, no dedicated extraction of κ4 from HH measurements is currently available."
→ It would be good to quote a rough order of magnitude of sensitivity of HH to help motivate HHH search.
L65: — (SPANET) has recently been demonstrated to be effective for both HH and HHH event reconstruction… You are citing theory studies, I dont think these are “demostrations”, these studies shown that SPANET is potentially effective.
L79: Is any offline b-tagging applied when training the inclusive classifier?
QCD MC for the multijet estimate in training.
For HH→4b the QCD MC is completely inadequate; I would assume the same holds for HHH→6b. How many total QCD MC events have exactly 6 medium b-tags? What is their average event weight? How many of these events pass the first classifier?
L92: The background reductions are quoted, which is useful, but it would also be good to quote the relative background composition. For example, is the background dominated by QCD multijet production?
L107: Are transfer factors used to scale from events failing the 6b/4b selection to events passing it? This seems necessary, but it is not explained how they are derived. Please clarify in the text.
L109: "discretized b and bb tagging scores are replaced with scores sampled from events in the SR, preserving inter-jet correlations."
What is meant by inter-jet correlations being preserved through score replacement? How is the replacement score chosen — is the inclusive distribution sampled randomly, or is event matching performed? For QCD, is there not a strong correlation between jet flavor and the underlying hard process? For example, ggg can produce 6b but ggq cannot. How is this handled?
L113: How large are these uncertainties?
L115: There is no comment on whether the prediction agrees in the tt̄ control region. This is odd, given that agreement in the inclusive distribution is mentioned in the preceding sentence. Does it not agree?
L151: "To capture residual discrepancies, the shape of the background model in the 1H category is assigned as a shape uncertainty in the HHH SR and treated as uncorrelated across categories."
It is not clear why this is a sensible approach. Why should a bias in the 3H region resemble the shape in the 1H region? This is certainly not the case in HH→4b, where the "1H" region (2b+2j) is a poor model for the "2H" region (2b+2b). Correcting for this difference is the main crux of that analysis.
L156: The relative background composition is still not clear to the reader at this point. Please clarify. Also clarify how dominant the statistical uncertainties are relative to systematics — for example, are systematics at the 30% level, 1%, or negligible? This should be stated explicitly.
L163: Can you comment on the quality of the pre-fit agreement, or provide the post-fit background-only χ²?
L169: SIngle H production modes…. Was this not already mention in paragraph L122-L126?
Figure 1: Where is the tt̄ contribution?
Please add control region distributions and pre-fit signal region distributions in the supplementary material. These are far more informative than eg: the CMS detector description paragraph.
L292-L298. What was this performed? Is this a geometrical matching between genpart and reco objects? What is the distance? If one ak8 jet is matched to two bquarks, and you also have two ak4 jets matched to two quarks in the same event, what was priority?
cheers john
My notes
HIG-24-012
This is great work. It is a shame it is only being made public in a letter, would be better in a longer PRD-style paper.
L15 "Although HH production is primarily sensitive to κ3, it also depends on κ4 through electroweak (EW) corrections; however, no dedicated extraction of κ4 from HH measurements is currently available."
→ It would be good to quote a rough order of magnitude of sensitivity of HH to help motivate HHH search.
(Note that σ(3H) ≈ 0.002 × σ(2H).)
L79: Is any offline b-tagging applied when training the inclusive classifier?
QCD MC for the multijet estimate in training.
For HH→4b the QCD MC is completely inadequate; I would assume the same holds for HHH→6b. How many total QCD MC events have exactly 6 medium b-tags? What is their average event weight? How many of these events pass the first classifier?
L92: The background reductions are quoted, which is useful, but it would also be good to quote the relative background composition. For example, is the background dominated by QCD multijet production?
L107: Are transfer factors used to scale from events failing the 6b/4b selection to events passing it? This seems necessary, but it is not explained how they are derived. Please clarify in the text.
L109: "discretized b and bb tagging scores are replaced with scores sampled from events in the SR, preserving inter-jet correlations."
What is meant by inter-jet correlations being preserved through score replacement? How is the replacement score chosen — is the inclusive distribution sampled randomly, or is event matching performed? For QCD, is there not a strong correlation between jet flavor and the underlying hard process? For example, ggg can produce 6b but ggq cannot. How is this handled?
L113: How large are these uncertainties?
L115: There is no comment on whether the prediction agrees in the tt̄ control region. This is odd, given that agreement in the inclusive distribution is mentioned in the preceding sentence. Does it not agree?
L151: "To capture residual discrepancies, the shape of the background model in the 1H category is assigned as a shape uncertainty in the HHH SR and treated as uncorrelated across categories."
It is not clear why this is a sensible approach. Why should a bias in the 3H region resemble the shape in the 1H region? This is certainly not the case in HH→4b, where the "1H" region (2b+2j) is a poor model for the "2H" region (2b+2b). Correcting for this difference is the main crux of that analysis.