Skip to contents

Complete Results

These results are based on Carter (2019) data-generating mechanism with a total of 756 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MMPH (default) 0.089 1 WAAPWLS (default) 0.105
2 WAAPWLS (default) 0.105 2 PEESE (default) 0.107
3 PEESE (default) 0.107 3 MMPH (default) 0.109
4 trimfill (default) 0.113 4 trimfill (default) 0.113
5 PETPEESE (default) 0.119 5 PETPEESE (default) 0.119
6 FMA (default) 0.124 6 FMA (default) 0.124
6 WLS (default) 0.124 6 WLS (default) 0.124
8 WILS (default) 0.128 8 WILS (default) 0.128
9 RoBMA (PSMA) 0.135 9 RoBMA (PSMA) 0.137
10 EK (default) 0.149 10 EK (default) 0.149
11 PET (default) 0.149 11 PET (default) 0.149
12 AK (AK1) 0.163 12 AK (AK1) 0.161
13 RMA (default) 0.164 13 RMA (default) 0.164
14 pcurve (default) 0.169 14 pcurve (default) 0.165
15 MAIVE (default) 0.170 15 MAIVE (default) 0.170
16 AK (AK2) 0.189 16 RTMA (relaxed) 0.217
17 MAN (default) 0.242 17 MAN (default) 0.220
18 mean (default) 0.249 18 AK (AK2) 0.221
19 SM (3PSM) 0.300 19 mean (default) 0.249
20 RTMA (relaxed) 0.343 20 SM (3PSM) 0.288
21 MAIVE (WAIVE) 0.378 21 MAIVE (WAIVE) 0.378
22 puniform (default) 0.453 22 puniform (default) 0.421
23 SM (4PSM) 0.459 23 SM (4PSM) 0.479
24 puniform (star) 169.671 24 puniform (star) 169.671

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 PETPEESE (default) 0.002 1 PETPEESE (default) 0.002
2 MMPH (default) 0.002 2 AK (AK1) 0.022
3 AK (AK1) 0.022 3 PEESE (default) 0.025
4 PEESE (default) 0.025 4 MMPH (default) 0.029
5 puniform (default) 0.028 5 puniform (default) 0.033
6 RTMA (relaxed) -0.047 6 AK (AK2) -0.035
7 pcurve (default) 0.052 7 pcurve (default) 0.048
8 WILS (default) -0.052 8 WILS (default) -0.052
9 EK (default) -0.052 9 EK (default) -0.052
10 PET (default) -0.052 10 PET (default) -0.052
11 WAAPWLS (default) 0.054 11 WAAPWLS (default) 0.054
12 MAIVE (default) -0.055 12 MAIVE (default) -0.055
13 trimfill (default) 0.067 13 trimfill (default) 0.067
14 AK (AK2) -0.075 14 RoBMA (PSMA) -0.088
15 FMA (default) 0.092 15 FMA (default) 0.092
15 WLS (default) 0.092 15 WLS (default) 0.092
17 RoBMA (PSMA) -0.094 17 MAN (default) -0.094
18 SM (3PSM) -0.108 18 SM (3PSM) -0.100
19 RMA (default) 0.150 19 RTMA (relaxed) 0.121
20 SM (4PSM) -0.200 20 RMA (default) 0.150
21 MAN (default) -0.212 21 SM (4PSM) -0.203
22 mean (default) 0.235 22 mean (default) 0.235
23 MAIVE (WAIVE) -0.260 23 MAIVE (WAIVE) -0.260
24 puniform (star) -31.860 24 puniform (star) -31.860

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RMA (default) 0.041 1 RMA (default) 0.041
2 trimfill (default) 0.045 2 trimfill (default) 0.045
3 mean (default) 0.052 3 mean (default) 0.052
4 MMPH (default) 0.054 4 FMA (default) 0.058
5 FMA (default) 0.058 4 WLS (default) 0.058
5 WLS (default) 0.058 6 MMPH (default) 0.063
7 MAN (default) 0.062 7 WAAPWLS (default) 0.072
8 WAAPWLS (default) 0.072 8 PEESE (default) 0.074
9 RoBMA (PSMA) 0.074 9 RoBMA (PSMA) 0.086
10 PEESE (default) 0.074 10 WILS (default) 0.087
11 WILS (default) 0.087 11 MAN (default) 0.088
12 PETPEESE (default) 0.096 12 pcurve (default) 0.094
13 pcurve (default) 0.097 13 PETPEESE (default) 0.096
14 PET (default) 0.119 14 AK (AK1) 0.117
15 EK (default) 0.119 15 PET (default) 0.119
16 AK (AK1) 0.119 16 EK (default) 0.119
17 MAIVE (default) 0.120 17 MAIVE (default) 0.120
18 AK (AK2) 0.130 18 RTMA (relaxed) 0.131
19 RTMA (relaxed) 0.210 19 AK (AK2) 0.184
20 MAIVE (WAIVE) 0.223 20 MAIVE (WAIVE) 0.223
21 SM (3PSM) 0.246 21 SM (3PSM) 0.235
22 SM (4PSM) 0.366 22 puniform (default) 0.348
23 puniform (default) 0.381 23 SM (4PSM) 0.386
24 puniform (star) 165.831 24 puniform (star) 165.831

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MMPH (default) 0.931 1 puniform (star) 1.402
2 puniform (star) 1.402 2 WAAPWLS (default) 1.473
3 WAAPWLS (default) 1.473 3 MMPH (default) 1.618
4 PETPEESE (default) 1.679 4 PETPEESE (default) 1.679
5 RoBMA (PSMA) 1.721 5 PEESE (default) 1.725
6 PEESE (default) 1.725 6 RoBMA (PSMA) 1.761
7 EK (default) 1.827 7 SM (3PSM) 1.824
8 SM (3PSM) 1.841 8 EK (default) 1.827
9 PET (default) 1.901 9 PET (default) 1.901
10 MAIVE (default) 2.210 10 MAIVE (default) 2.210
11 WILS (default) 2.220 11 WILS (default) 2.220
12 trimfill (default) 2.241 12 trimfill (default) 2.244
13 puniform (default) 2.587 13 puniform (default) 2.533
14 WLS (default) 2.680 14 WLS (default) 2.680
15 SM (4PSM) 2.879 15 SM (4PSM) 2.820
16 AK (AK1) 2.896 16 AK (AK1) 2.822
17 FMA (default) 3.172 17 FMA (default) 3.172
18 AK (AK2) 3.599 18 AK (AK2) 3.480
19 RMA (default) 4.078 19 RMA (default) 4.078
20 MAIVE (WAIVE) 4.682 20 RTMA (relaxed) 4.102
21 MAN (default) 4.970 21 MAIVE (WAIVE) 4.682
22 RTMA (relaxed) 5.648 22 MAN (default) 4.996
23 mean (default) 7.200 23 mean (default) 7.200
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RTMA (relaxed) 0.728 1 AK (AK2) 0.674
2 AK (AK2) 0.680 2 RoBMA (PSMA) 0.645
3 MMPH (default) 0.679 3 puniform (star) 0.643
4 RoBMA (PSMA) 0.647 4 SM (3PSM) 0.630
5 puniform (star) 0.643 5 SM (4PSM) 0.629
6 SM (3PSM) 0.635 6 WAAPWLS (default) 0.604
7 SM (4PSM) 0.634 7 MMPH (default) 0.591
8 WAAPWLS (default) 0.604 8 AK (AK1) 0.589
9 AK (AK1) 0.590 9 MAIVE (default) 0.563
10 MAIVE (default) 0.563 10 puniform (default) 0.549
11 puniform (default) 0.549 11 PETPEESE (default) 0.512
12 trimfill (default) 0.512 12 trimfill (default) 0.512
13 PETPEESE (default) 0.512 13 EK (default) 0.477
14 EK (default) 0.477 14 MAIVE (WAIVE) 0.468
15 MAIVE (WAIVE) 0.468 15 PEESE (default) 0.464
16 PEESE (default) 0.464 16 PET (default) 0.463
17 PET (default) 0.463 17 RTMA (relaxed) 0.439
18 WILS (default) 0.415 18 WILS (default) 0.415
19 WLS (default) 0.397 19 WLS (default) 0.397
20 RMA (default) 0.359 20 RMA (default) 0.359
21 MAN (default) 0.327 21 FMA (default) 0.307
22 FMA (default) 0.307 22 MAN (default) 0.275
23 mean (default) 0.179 23 mean (default) 0.179
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.091 1 FMA (default) 0.091
2 WLS (default) 0.133 2 WLS (default) 0.133
3 mean (default) 0.148 3 mean (default) 0.148
4 PEESE (default) 0.153 4 PEESE (default) 0.153
5 trimfill (default) 0.160 5 trimfill (default) 0.160
6 WILS (default) 0.161 6 WILS (default) 0.161
7 RMA (default) 0.163 7 RMA (default) 0.163
8 WAAPWLS (default) 0.189 8 WAAPWLS (default) 0.189
9 PETPEESE (default) 0.190 9 PETPEESE (default) 0.190
10 MMPH (default) 0.233 10 MAN (default) 0.215
11 PET (default) 0.244 11 MMPH (default) 0.226
12 EK (default) 0.264 12 PET (default) 0.244
13 RoBMA (PSMA) 0.279 13 EK (default) 0.264
14 MAN (default) 0.302 14 RoBMA (PSMA) 0.277
15 puniform (star) 0.341 15 puniform (star) 0.341
16 MAIVE (default) 0.381 16 MAIVE (default) 0.381
17 puniform (default) 0.540 17 puniform (default) 0.494
18 SM (3PSM) 0.615 18 SM (3PSM) 0.557
19 MAIVE (WAIVE) 0.638 19 RTMA (relaxed) 0.630
20 SM (4PSM) 1.079 20 MAIVE (WAIVE) 0.638
21 AK (AK1) 1.787 21 SM (4PSM) 0.985
22 AK (AK2) 2.464 22 AK (AK1) 1.711
23 RTMA (relaxed) 3.758 23 AK (AK2) 2.604
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 2.323 1 RoBMA (PSMA) 2.252
2 puniform (default) 1.987 2 puniform (default) 1.979
3 MMPH (default) 1.460 3 AK (AK2) 1.165
4 AK (AK2) 1.351 4 MMPH (default) 1.117
5 PETPEESE (default) 1.091 5 PETPEESE (default) 1.091
6 AK (AK1) 1.030 6 AK (AK1) 1.027
7 MAIVE (default) 0.973 7 MAIVE (default) 0.973
8 PET (default) 0.961 8 PET (default) 0.961
9 EK (default) 0.960 9 EK (default) 0.960
10 SM (3PSM) 0.909 10 SM (3PSM) 0.895
11 WAAPWLS (default) 0.821 11 WAAPWLS (default) 0.821
12 puniform (star) 0.798 12 puniform (star) 0.798
13 trimfill (default) 0.793 13 trimfill (default) 0.793
14 RMA (default) 0.764 14 RTMA (relaxed) 0.770
15 PEESE (default) 0.627 15 RMA (default) 0.764
16 WLS (default) 0.621 16 PEESE (default) 0.627
17 WILS (default) 0.573 17 WLS (default) 0.621
18 RTMA (relaxed) 0.526 18 WILS (default) 0.573
19 FMA (default) 0.385 19 MAN (default) 0.460
20 SM (4PSM) 0.368 20 FMA (default) 0.385
21 mean (default) 0.351 21 SM (4PSM) 0.381
22 MAN (default) 0.289 22 mean (default) 0.351
23 MAIVE (WAIVE) 0.073 23 MAIVE (WAIVE) 0.073
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 PETPEESE (default) -4.367 1 PETPEESE (default) -4.367
2 PET (default) -4.242 2 PET (default) -4.242
3 EK (default) -4.242 3 EK (default) -4.242
4 WAAPWLS (default) -4.199 4 WAAPWLS (default) -4.199
5 PEESE (default) -3.861 5 MMPH (default) -4.012
6 MMPH (default) -3.809 6 PEESE (default) -3.861
7 puniform (default) -3.695 7 AK (AK2) -3.721
8 WLS (default) -3.403 8 puniform (default) -3.695
9 trimfill (default) -3.240 9 WLS (default) -3.403
10 AK (AK1) -3.054 10 trimfill (default) -3.241
11 FMA (default) -3.028 11 AK (AK1) -3.057
12 RMA (default) -2.866 12 FMA (default) -3.028
13 SM (3PSM) -2.688 13 RMA (default) -2.866
14 puniform (star) -2.427 14 SM (3PSM) -2.705
15 MAIVE (default) -2.347 15 puniform (star) -2.427
16 WILS (default) -2.292 16 MAIVE (default) -2.347
17 RoBMA (PSMA) -2.270 17 RTMA (relaxed) -2.307
18 mean (default) -2.164 18 WILS (default) -2.292
19 AK (AK2) -1.915 19 RoBMA (PSMA) -2.266
20 SM (4PSM) -1.222 20 mean (default) -2.164
21 RTMA (relaxed) -0.991 21 MAN (default) -1.485
22 MAN (default) -0.498 22 SM (4PSM) -1.314
23 MAIVE (WAIVE) 0.007 23 MAIVE (WAIVE) 0.007
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK2) 0.276 1 RoBMA (PSMA) 0.283
2 RoBMA (PSMA) 0.278 2 MAIVE (WAIVE) 0.288
3 MAIVE (WAIVE) 0.288 3 MAIVE (default) 0.344
4 MAIVE (default) 0.344 4 PETPEESE (default) 0.403
5 MMPH (default) 0.374 5 PET (default) 0.421
6 PETPEESE (default) 0.403 6 EK (default) 0.421
7 PET (default) 0.421 7 AK (AK2) 0.439
8 EK (default) 0.421 8 puniform (star) 0.443
9 puniform (star) 0.443 9 puniform (default) 0.457
10 puniform (default) 0.457 10 SM (3PSM) 0.473
11 RTMA (relaxed) 0.462 11 MMPH (default) 0.477
12 SM (3PSM) 0.468 12 WAAPWLS (default) 0.506
13 WAAPWLS (default) 0.506 13 SM (4PSM) 0.545
14 MAN (default) 0.540 14 WILS (default) 0.557
15 SM (4PSM) 0.543 15 MAN (default) 0.618
16 WILS (default) 0.557 16 PEESE (default) 0.638
17 PEESE (default) 0.638 17 RTMA (relaxed) 0.640
18 AK (AK1) 0.642 18 AK (AK1) 0.642
19 trimfill (default) 0.660 19 trimfill (default) 0.660
20 WLS (default) 0.690 20 WLS (default) 0.690
21 RMA (default) 0.702 21 RMA (default) 0.702
22 FMA (default) 0.796 22 FMA (default) 0.796
23 mean (default) 0.827 23 mean (default) 0.827
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 mean (default) 0.995 1 mean (default) 0.995
2 FMA (default) 0.994 2 FMA (default) 0.994
3 RMA (default) 0.989 3 RMA (default) 0.989
4 WLS (default) 0.985 4 WLS (default) 0.985
5 trimfill (default) 0.980 5 trimfill (default) 0.980
6 AK (AK1) 0.979 6 AK (AK1) 0.979
7 PEESE (default) 0.963 7 PEESE (default) 0.963
8 WAAPWLS (default) 0.925 8 RTMA (relaxed) 0.939
9 puniform (default) 0.913 9 WAAPWLS (default) 0.925
10 MMPH (default) 0.906 10 puniform (default) 0.913
11 PETPEESE (default) 0.899 11 MMPH (default) 0.913
12 EK (default) 0.885 12 PETPEESE (default) 0.899
13 PET (default) 0.885 13 EK (default) 0.885
14 WILS (default) 0.846 14 PET (default) 0.885
15 SM (3PSM) 0.806 15 AK (AK2) 0.858
16 AK (AK2) 0.748 16 WILS (default) 0.846
17 puniform (star) 0.746 17 MAN (default) 0.817
18 MAIVE (default) 0.706 18 SM (3PSM) 0.812
19 SM (4PSM) 0.697 19 puniform (star) 0.746
20 MAN (default) 0.694 20 SM (4PSM) 0.709
21 RoBMA (PSMA) 0.642 21 MAIVE (default) 0.706
22 RTMA (relaxed) 0.614 22 RoBMA (PSMA) 0.646
23 MAIVE (WAIVE) 0.281 23 MAIVE (WAIVE) 0.281
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: No Questionable Research Practices

These results are based on Carter (2019) data-generating mechanism with a total of 252 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RMA (default) 0.057 1 RMA (default) 0.057
2 FMA (default) 0.063 2 FMA (default) 0.063
2 WLS (default) 0.063 2 WLS (default) 0.063
4 trimfill (default) 0.063 4 trimfill (default) 0.063
5 WAAPWLS (default) 0.077 5 WAAPWLS (default) 0.077
6 MMPH (default) 0.079 6 MMPH (default) 0.079
7 PEESE (default) 0.094 7 PEESE (default) 0.094
8 mean (default) 0.104 8 mean (default) 0.104
9 RoBMA (PSMA) 0.109 9 RoBMA (PSMA) 0.108
10 PETPEESE (default) 0.118 10 PETPEESE (default) 0.118
11 WILS (default) 0.131 11 WILS (default) 0.131
12 SM (3PSM) 0.134 12 RTMA (relaxed) 0.133
13 EK (default) 0.161 13 SM (3PSM) 0.134
14 PET (default) 0.161 14 EK (default) 0.161
15 MAIVE (default) 0.181 15 PET (default) 0.161
16 SM (4PSM) 0.196 16 MAIVE (default) 0.181
17 RTMA (relaxed) 0.206 17 SM (4PSM) 0.198
18 pcurve (default) 0.235 18 pcurve (default) 0.222
19 AK (AK2) 0.246 19 MAN (default) 0.241
20 MAN (default) 0.274 20 AK (AK1) 0.289
21 AK (AK1) 0.297 21 AK (AK2) 0.356
22 MAIVE (WAIVE) 0.388 22 MAIVE (WAIVE) 0.388
23 puniform (default) 0.952 23 puniform (default) 0.862
24 puniform (star) 46.988 24 puniform (star) 46.988

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.004 1 puniform (default) -0.001
1 WLS (default) 0.004 2 FMA (default) 0.004
3 WAAPWLS (default) -0.007 2 WLS (default) 0.004
4 puniform (default) -0.014 4 WAAPWLS (default) -0.007
5 trimfill (default) -0.025 5 trimfill (default) -0.025
6 RMA (default) 0.028 6 RMA (default) 0.028
7 AK (AK1) -0.031 7 AK (AK1) -0.032
8 PEESE (default) -0.038 8 PEESE (default) -0.038
9 PETPEESE (default) -0.052 9 RTMA (relaxed) -0.039
10 MMPH (default) -0.058 10 MMPH (default) -0.048
11 EK (default) -0.073 11 PETPEESE (default) -0.052
12 PET (default) -0.073 12 pcurve (default) 0.063
13 pcurve (default) 0.075 13 EK (default) -0.073
14 mean (default) 0.075 14 PET (default) -0.073
15 SM (3PSM) -0.085 15 AK (AK2) -0.074
16 RoBMA (PSMA) -0.088 16 mean (default) 0.075
17 WILS (default) -0.096 17 SM (3PSM) -0.085
18 AK (AK2) -0.101 18 RoBMA (PSMA) -0.087
19 SM (4PSM) -0.106 19 WILS (default) -0.096
20 MAIVE (default) -0.113 20 SM (4PSM) -0.106
21 RTMA (relaxed) -0.128 21 MAIVE (default) -0.113
22 MAN (default) -0.267 22 MAN (default) -0.210
23 MAIVE (WAIVE) -0.277 23 MAIVE (WAIVE) -0.277
24 puniform (star) -4.095 24 puniform (star) -4.095

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RMA (default) 0.043 1 RMA (default) 0.043
2 MAN (default) 0.046 2 trimfill (default) 0.048
3 trimfill (default) 0.048 3 MMPH (default) 0.050
4 MMPH (default) 0.049 4 mean (default) 0.055
5 RoBMA (PSMA) 0.054 5 RoBMA (PSMA) 0.057
6 mean (default) 0.055 6 FMA (default) 0.058
7 FMA (default) 0.058 6 WLS (default) 0.058
7 WLS (default) 0.058 8 MAN (default) 0.066
9 WAAPWLS (default) 0.074 9 WAAPWLS (default) 0.074
10 WILS (default) 0.076 10 WILS (default) 0.076
11 SM (3PSM) 0.081 11 SM (3PSM) 0.080
12 PEESE (default) 0.082 12 PEESE (default) 0.082
13 PETPEESE (default) 0.101 13 PETPEESE (default) 0.101
14 RTMA (relaxed) 0.122 14 RTMA (relaxed) 0.108
15 MAIVE (default) 0.122 15 MAIVE (default) 0.122
16 EK (default) 0.135 16 EK (default) 0.135
17 PET (default) 0.135 17 PET (default) 0.135
18 SM (4PSM) 0.136 18 SM (4PSM) 0.137
19 pcurve (default) 0.164 19 pcurve (default) 0.156
20 AK (AK2) 0.187 20 MAIVE (WAIVE) 0.234
21 MAIVE (WAIVE) 0.234 21 AK (AK1) 0.268
22 AK (AK1) 0.275 22 AK (AK2) 0.310
23 puniform (default) 0.888 23 puniform (default) 0.800
24 puniform (star) 46.707 24 puniform (star) 46.707

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RMA (default) 0.419 1 RMA (default) 0.419
2 WLS (default) 0.511 2 WLS (default) 0.511
3 trimfill (default) 0.533 3 trimfill (default) 0.533
4 WAAPWLS (default) 0.539 4 WAAPWLS (default) 0.539
5 MMPH (default) 0.677 5 MMPH (default) 0.700
6 FMA (default) 0.863 6 FMA (default) 0.863
7 PEESE (default) 0.977 7 PEESE (default) 0.977
8 puniform (star) 1.130 8 RTMA (relaxed) 1.085
9 PETPEESE (default) 1.153 9 puniform (star) 1.130
10 RoBMA (PSMA) 1.385 10 PETPEESE (default) 1.153
11 SM (3PSM) 1.563 11 RoBMA (PSMA) 1.363
12 EK (default) 1.711 12 SM (3PSM) 1.559
13 mean (default) 1.735 13 EK (default) 1.711
14 SM (4PSM) 1.762 14 SM (4PSM) 1.717
15 PET (default) 1.784 15 mean (default) 1.735
16 MAIVE (default) 1.918 16 PET (default) 1.784
17 RTMA (relaxed) 2.320 17 MAIVE (default) 1.918
18 WILS (default) 2.636 18 puniform (default) 2.538
19 puniform (default) 2.690 19 WILS (default) 2.636
20 MAIVE (WAIVE) 3.978 20 MAIVE (WAIVE) 3.978
21 AK (AK1) 5.500 21 AK (AK1) 5.273
22 AK (AK2) 5.978 22 MAN (default) 6.378
23 MAN (default) 7.175 23 AK (AK2) 6.407
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RTMA (relaxed) 0.831 1 WAAPWLS (default) 0.798
2 WAAPWLS (default) 0.798 2 RTMA (relaxed) 0.791
3 trimfill (default) 0.749 3 trimfill (default) 0.748
4 RMA (default) 0.736 4 RMA (default) 0.736
5 AK (AK1) 0.711 5 AK (AK1) 0.711
6 WLS (default) 0.698 6 WLS (default) 0.698
7 MMPH (default) 0.672 7 AK (AK2) 0.659
8 puniform (star) 0.656 8 puniform (star) 0.656
9 AK (AK2) 0.653 9 MMPH (default) 0.655
10 MAIVE (default) 0.651 10 MAIVE (default) 0.651
11 RoBMA (PSMA) 0.639 11 RoBMA (PSMA) 0.640
12 SM (4PSM) 0.619 12 SM (4PSM) 0.619
13 PEESE (default) 0.618 13 PEESE (default) 0.618
14 PETPEESE (default) 0.615 14 PETPEESE (default) 0.615
15 puniform (default) 0.606 15 puniform (default) 0.607
16 SM (3PSM) 0.595 16 SM (3PSM) 0.595
17 EK (default) 0.568 17 EK (default) 0.568
18 PET (default) 0.554 18 PET (default) 0.554
19 MAIVE (WAIVE) 0.541 19 MAIVE (WAIVE) 0.541
20 FMA (default) 0.530 20 FMA (default) 0.530
21 mean (default) 0.437 21 mean (default) 0.437
22 WILS (default) 0.402 22 WILS (default) 0.402
23 MAN (default) 0.169 23 MAN (default) 0.248
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.096 1 FMA (default) 0.096
2 WILS (default) 0.138 2 WILS (default) 0.138
3 WLS (default) 0.143 3 WLS (default) 0.143
4 mean (default) 0.149 4 mean (default) 0.149
5 trimfill (default) 0.175 5 trimfill (default) 0.175
6 RMA (default) 0.177 6 RMA (default) 0.177
7 PEESE (default) 0.184 7 PEESE (default) 0.184
8 WAAPWLS (default) 0.206 8 MAN (default) 0.190
9 RoBMA (PSMA) 0.211 9 WAAPWLS (default) 0.206
10 MAN (default) 0.211 10 MMPH (default) 0.210
11 MMPH (default) 0.214 11 RoBMA (PSMA) 0.210
12 PETPEESE (default) 0.237 12 PETPEESE (default) 0.237
13 SM (3PSM) 0.268 13 SM (3PSM) 0.264
14 puniform (star) 0.270 14 puniform (star) 0.270
15 PET (default) 0.296 15 PET (default) 0.296
16 EK (default) 0.322 16 EK (default) 0.322
17 SM (4PSM) 0.455 17 SM (4PSM) 0.410
18 MAIVE (default) 0.477 18 MAIVE (default) 0.477
19 MAIVE (WAIVE) 0.762 19 RTMA (relaxed) 0.599
20 puniform (default) 0.976 20 MAIVE (WAIVE) 0.762
21 RTMA (relaxed) 1.751 21 puniform (default) 0.845
22 AK (AK2) 4.827 22 AK (AK1) 4.810
23 AK (AK1) 5.038 23 AK (AK2) 5.426
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RoBMA (PSMA) 3.582 1 RoBMA (PSMA) 3.669
2 AK (AK1) 2.778 2 AK (AK1) 2.770
3 RMA (default) 2.193 3 RMA (default) 2.193
4 trimfill (default) 2.152 4 trimfill (default) 2.152
5 puniform (default) 1.977 5 puniform (default) 1.961
6 MMPH (default) 1.682 6 RTMA (relaxed) 1.931
7 WLS (default) 1.676 7 WLS (default) 1.676
8 WAAPWLS (default) 1.563 8 MMPH (default) 1.668
9 RTMA (relaxed) 1.553 9 WAAPWLS (default) 1.563
10 puniform (star) 1.441 10 AK (AK2) 1.556
11 PEESE (default) 1.327 11 puniform (star) 1.441
12 PETPEESE (default) 1.272 12 PEESE (default) 1.327
13 AK (AK2) 1.240 13 PETPEESE (default) 1.272
14 SM (3PSM) 1.112 14 SM (3PSM) 1.120
15 FMA (default) 1.095 15 FMA (default) 1.095
16 EK (default) 1.066 16 EK (default) 1.066
16 PET (default) 1.066 16 PET (default) 1.066
18 mean (default) 1.033 18 mean (default) 1.033
19 MAIVE (default) 0.980 19 MAIVE (default) 0.980
20 SM (4PSM) 0.893 20 SM (4PSM) 0.915
21 WILS (default) 0.856 21 WILS (default) 0.856
22 MAN (default) 0.166 22 MAN (default) 0.379
23 MAIVE (WAIVE) -0.131 23 MAIVE (WAIVE) -0.131
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 RMA (default) -6.535 1 RMA (default) -6.535
2 AK (AK1) -6.401 2 AK (AK1) -6.404
3 WLS (default) -6.229 3 WLS (default) -6.229
4 trimfill (default) -6.222 4 trimfill (default) -6.224
5 FMA (default) -6.193 5 FMA (default) -6.193
6 mean (default) -5.430 6 mean (default) -5.430
7 PEESE (default) -5.202 7 PEESE (default) -5.202
8 WAAPWLS (default) -4.263 8 MMPH (default) -4.914
9 PETPEESE (default) -4.199 9 WAAPWLS (default) -4.263
10 MMPH (default) -4.183 10 PETPEESE (default) -4.199
11 puniform (default) -4.028 11 puniform (default) -4.030
12 EK (default) -3.956 12 EK (default) -3.956
12 PET (default) -3.956 12 PET (default) -3.956
14 puniform (star) -3.690 14 RTMA (relaxed) -3.827
15 RoBMA (PSMA) -3.314 15 puniform (star) -3.690
16 SM (3PSM) -3.252 16 AK (AK2) -3.483
17 WILS (default) -3.212 17 RoBMA (PSMA) -3.326
18 RTMA (relaxed) -3.033 18 SM (3PSM) -3.259
19 SM (4PSM) -2.447 19 WILS (default) -3.212
20 MAIVE (default) -2.288 20 SM (4PSM) -2.519
21 AK (AK2) -1.757 21 MAIVE (default) -2.288
22 MAN (default) -0.327 22 MAN (default) -0.800
23 MAIVE (WAIVE) -0.127 23 MAIVE (WAIVE) -0.127
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK1) 0.114 1 AK (AK1) 0.114
2 RoBMA (PSMA) 0.143 2 RoBMA (PSMA) 0.143
3 trimfill (default) 0.155 3 trimfill (default) 0.155
4 RMA (default) 0.186 4 RMA (default) 0.186
5 WAAPWLS (default) 0.220 5 WAAPWLS (default) 0.220
6 WLS (default) 0.227 6 WLS (default) 0.227
7 RTMA (relaxed) 0.244 7 RTMA (relaxed) 0.239
8 MAIVE (WAIVE) 0.276 8 MAIVE (WAIVE) 0.276
9 MMPH (default) 0.288 9 MMPH (default) 0.291
10 MAIVE (default) 0.293 10 MAIVE (default) 0.293
11 PETPEESE (default) 0.302 11 PETPEESE (default) 0.302
12 AK (AK2) 0.312 12 AK (AK2) 0.307
13 PEESE (default) 0.312 13 PEESE (default) 0.312
14 puniform (star) 0.319 14 puniform (star) 0.319
15 EK (default) 0.348 15 EK (default) 0.348
15 PET (default) 0.348 15 PET (default) 0.348
17 puniform (default) 0.378 17 puniform (default) 0.377
18 SM (4PSM) 0.413 18 SM (4PSM) 0.413
19 SM (3PSM) 0.428 19 SM (3PSM) 0.428
20 FMA (default) 0.442 20 FMA (default) 0.442
21 mean (default) 0.499 21 mean (default) 0.499
22 WILS (default) 0.501 22 WILS (default) 0.501
23 MAN (default) 0.584 23 MAN (default) 0.584
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 mean (default) 0.986 1 mean (default) 0.986
2 FMA (default) 0.985 2 FMA (default) 0.985
3 RMA (default) 0.969 3 RMA (default) 0.969
4 WLS (default) 0.964 4 WLS (default) 0.964
5 trimfill (default) 0.949 5 trimfill (default) 0.949
6 AK (AK1) 0.949 6 AK (AK1) 0.948
7 PEESE (default) 0.915 7 PEESE (default) 0.915
8 puniform (default) 0.885 8 puniform (default) 0.885
9 MMPH (default) 0.876 9 MMPH (default) 0.879
10 WILS (default) 0.876 10 WILS (default) 0.876
11 WAAPWLS (default) 0.858 11 RTMA (relaxed) 0.863
12 PETPEESE (default) 0.845 12 WAAPWLS (default) 0.858
13 EK (default) 0.831 13 PETPEESE (default) 0.845
13 PET (default) 0.831 14 EK (default) 0.831
15 SM (3PSM) 0.819 14 PET (default) 0.831
16 puniform (star) 0.811 16 SM (3PSM) 0.821
17 SM (4PSM) 0.732 17 puniform (star) 0.811
18 AK (AK2) 0.711 18 AK (AK2) 0.795
19 RTMA (relaxed) 0.698 19 MAN (default) 0.749
20 MAN (default) 0.674 20 SM (4PSM) 0.739
21 RoBMA (PSMA) 0.674 21 RoBMA (PSMA) 0.676
22 MAIVE (default) 0.649 22 MAIVE (default) 0.649
23 MAIVE (WAIVE) 0.286 23 MAIVE (WAIVE) 0.286
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: Medium Questionable Research Practices

These results are based on Carter (2019) data-generating mechanism with a total of 252 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MMPH (default) 0.084 1 AK (AK1) 0.087
2 AK (AK1) 0.087 2 PEESE (default) 0.101
3 PEESE (default) 0.101 3 AK (AK2) 0.105
4 WAAPWLS (default) 0.107 4 WAAPWLS (default) 0.107
5 PETPEESE (default) 0.111 5 MMPH (default) 0.108
6 trimfill (default) 0.117 6 PETPEESE (default) 0.111
7 WILS (default) 0.122 7 trimfill (default) 0.117
8 FMA (default) 0.133 8 WILS (default) 0.122
8 WLS (default) 0.133 9 FMA (default) 0.133
10 RoBMA (PSMA) 0.135 9 WLS (default) 0.133
11 AK (AK2) 0.136 11 RoBMA (PSMA) 0.137
12 pcurve (default) 0.143 12 pcurve (default) 0.142
13 EK (default) 0.145 13 EK (default) 0.145
14 PET (default) 0.145 14 PET (default) 0.145
15 MAIVE (default) 0.162 15 MAIVE (default) 0.162
16 RMA (default) 0.191 16 RMA (default) 0.191
17 MAN (default) 0.229 17 MAN (default) 0.204
18 puniform (default) 0.259 18 RTMA (relaxed) 0.222
19 SM (3PSM) 0.272 19 puniform (default) 0.251
20 mean (default) 0.284 20 SM (3PSM) 0.263
21 RTMA (relaxed) 0.329 21 mean (default) 0.284
22 MAIVE (WAIVE) 0.370 22 MAIVE (WAIVE) 0.370
23 SM (4PSM) 0.444 23 SM (4PSM) 0.462
24 puniform (star) 147.205 24 puniform (star) 147.205

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 PETPEESE (default) 0.011 1 PETPEESE (default) 0.011
2 MMPH (default) 0.023 2 AK (AK2) -0.018
3 RTMA (relaxed) -0.027 3 AK (AK1) 0.036
4 AK (AK1) 0.036 4 PEESE (default) 0.037
5 PEESE (default) 0.037 5 MAIVE (default) -0.041
6 MAIVE (default) -0.041 6 pcurve (default) 0.042
7 pcurve (default) 0.042 7 puniform (default) 0.046
8 puniform (default) 0.045 8 WILS (default) -0.048
9 WILS (default) -0.048 9 MMPH (default) 0.051
10 EK (default) -0.052 10 EK (default) -0.052
11 PET (default) -0.052 11 PET (default) -0.052
12 WAAPWLS (default) 0.068 12 WAAPWLS (default) 0.068
13 AK (AK2) -0.086 13 MAN (default) -0.078
14 SM (3PSM) -0.090 14 SM (3PSM) -0.084
15 trimfill (default) 0.090 15 RoBMA (PSMA) -0.085
16 RoBMA (PSMA) -0.091 16 trimfill (default) 0.090
17 FMA (default) 0.112 17 FMA (default) 0.112
17 WLS (default) 0.112 17 WLS (default) 0.112
19 RMA (default) 0.183 19 RTMA (relaxed) 0.161
20 SM (4PSM) -0.202 20 RMA (default) 0.183
21 MAN (default) -0.202 21 SM (4PSM) -0.204
22 MAIVE (WAIVE) -0.254 22 MAIVE (WAIVE) -0.254
23 mean (default) 0.277 23 mean (default) 0.277
24 puniform (star) -23.968 24 puniform (star) -23.968

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RMA (default) 0.041 1 RMA (default) 0.041
2 AK (AK1) 0.042 2 AK (AK1) 0.042
3 trimfill (default) 0.044 3 trimfill (default) 0.045
4 mean (default) 0.052 4 mean (default) 0.052
5 MMPH (default) 0.055 5 FMA (default) 0.059
6 FMA (default) 0.059 5 WLS (default) 0.059
6 WLS (default) 0.059 7 MMPH (default) 0.064
8 MAN (default) 0.062 8 pcurve (default) 0.070
9 pcurve (default) 0.071 9 PEESE (default) 0.074
10 PEESE (default) 0.074 10 WAAPWLS (default) 0.074
11 WAAPWLS (default) 0.074 11 AK (AK2) 0.081
12 RoBMA (PSMA) 0.078 12 WILS (default) 0.089
13 AK (AK2) 0.079 13 RoBMA (PSMA) 0.091
14 WILS (default) 0.089 14 MAN (default) 0.094
15 PETPEESE (default) 0.095 15 PETPEESE (default) 0.095
16 PET (default) 0.118 16 PET (default) 0.118
17 EK (default) 0.118 17 EK (default) 0.118
18 MAIVE (default) 0.118 18 MAIVE (default) 0.118
19 puniform (default) 0.183 19 RTMA (relaxed) 0.137
20 RTMA (relaxed) 0.207 20 puniform (default) 0.175
21 SM (3PSM) 0.220 21 SM (3PSM) 0.212
22 MAIVE (WAIVE) 0.221 22 MAIVE (WAIVE) 0.221
23 SM (4PSM) 0.343 23 SM (4PSM) 0.361
24 puniform (star) 144.875 24 puniform (star) 144.875

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MMPH (default) 0.823 1 AK (AK2) 0.960
2 AK (AK2) 1.212 2 AK (AK1) 1.240
3 AK (AK1) 1.235 3 puniform (star) 1.358
4 puniform (star) 1.358 4 WAAPWLS (default) 1.375
5 WAAPWLS (default) 1.375 5 PETPEESE (default) 1.437
6 PETPEESE (default) 1.437 6 PEESE (default) 1.462
7 PEESE (default) 1.462 7 MMPH (default) 1.599
8 SM (3PSM) 1.698 8 SM (3PSM) 1.674
9 RoBMA (PSMA) 1.724 9 EK (default) 1.745
10 EK (default) 1.745 10 RoBMA (PSMA) 1.780
11 PET (default) 1.819 11 PET (default) 1.819
12 WILS (default) 1.929 12 WILS (default) 1.929
13 MAIVE (default) 1.952 13 MAIVE (default) 1.952
14 trimfill (default) 2.200 14 trimfill (default) 2.202
15 puniform (default) 2.525 15 puniform (default) 2.516
16 WLS (default) 2.871 16 WLS (default) 2.871
17 SM (4PSM) 3.078 17 SM (4PSM) 2.975
18 FMA (default) 3.472 18 FMA (default) 3.472
19 MAN (default) 4.291 19 MAN (default) 3.951
20 MAIVE (WAIVE) 4.466 20 RTMA (relaxed) 4.061
21 RMA (default) 4.763 21 MAIVE (WAIVE) 4.466
22 RTMA (relaxed) 5.296 22 RMA (default) 4.763
23 mean (default) 8.442 23 mean (default) 8.442
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK2) 0.735 1 AK (AK2) 0.686
2 MMPH (default) 0.710 2 RoBMA (PSMA) 0.650
3 RTMA (relaxed) 0.708 3 SM (3PSM) 0.635
4 RoBMA (PSMA) 0.653 4 puniform (star) 0.633
5 SM (3PSM) 0.640 5 SM (4PSM) 0.614
6 puniform (star) 0.633 6 MMPH (default) 0.591
7 SM (4PSM) 0.619 7 WAAPWLS (default) 0.570
8 WAAPWLS (default) 0.570 8 MAIVE (default) 0.562
9 MAIVE (default) 0.562 9 AK (AK1) 0.551
10 AK (AK1) 0.553 10 puniform (default) 0.529
11 puniform (default) 0.529 11 PETPEESE (default) 0.512
12 PETPEESE (default) 0.512 12 MAIVE (WAIVE) 0.492
13 MAIVE (WAIVE) 0.492 13 EK (default) 0.468
14 EK (default) 0.468 14 PET (default) 0.453
15 PET (default) 0.453 15 PEESE (default) 0.446
16 PEESE (default) 0.446 16 trimfill (default) 0.426
17 trimfill (default) 0.427 17 WILS (default) 0.412
18 WILS (default) 0.412 18 RTMA (relaxed) 0.336
19 MAN (default) 0.347 19 MAN (default) 0.296
20 WLS (default) 0.275 20 WLS (default) 0.275
21 FMA (default) 0.213 21 FMA (default) 0.213
22 RMA (default) 0.191 22 RMA (default) 0.191
23 mean (default) 0.059 23 mean (default) 0.059
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.089 1 FMA (default) 0.089
2 WLS (default) 0.134 2 WLS (default) 0.134
3 mean (default) 0.147 3 mean (default) 0.147
4 PEESE (default) 0.150 4 PEESE (default) 0.150
5 trimfill (default) 0.160 5 trimfill (default) 0.160
6 WILS (default) 0.160 6 WILS (default) 0.160
7 RMA (default) 0.164 7 RMA (default) 0.164
8 AK (AK1) 0.168 8 AK (AK1) 0.168
9 PETPEESE (default) 0.184 9 PETPEESE (default) 0.184
10 WAAPWLS (default) 0.198 10 WAAPWLS (default) 0.198
11 PET (default) 0.237 11 MAN (default) 0.229
12 MMPH (default) 0.238 12 MMPH (default) 0.232
13 EK (default) 0.257 13 PET (default) 0.237
14 RoBMA (PSMA) 0.288 14 EK (default) 0.257
15 MAN (default) 0.307 15 AK (AK2) 0.266
16 puniform (star) 0.332 16 RoBMA (PSMA) 0.286
17 MAIVE (default) 0.371 17 puniform (star) 0.332
18 puniform (default) 0.375 18 puniform (default) 0.367
19 AK (AK2) 0.433 19 MAIVE (default) 0.371
20 SM (3PSM) 0.543 20 SM (3PSM) 0.498
21 MAIVE (WAIVE) 0.630 21 MAIVE (WAIVE) 0.630
22 SM (4PSM) 1.069 22 RTMA (relaxed) 0.724
23 RTMA (relaxed) 3.726 23 SM (4PSM) 0.939
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 puniform (default) 1.991 1 puniform (default) 1.983
2 RoBMA (PSMA) 1.807 2 RoBMA (PSMA) 1.743
3 AK (AK2) 1.595 3 PETPEESE (default) 1.261
4 MMPH (default) 1.518 4 MAIVE (default) 1.059
5 PETPEESE (default) 1.261 5 MMPH (default) 1.054
6 MAIVE (default) 1.059 6 PET (default) 1.012
7 PET (default) 1.012 7 EK (default) 1.012
8 EK (default) 1.012 8 AK (AK2) 0.847
9 SM (3PSM) 0.833 9 SM (3PSM) 0.828
10 WAAPWLS (default) 0.665 10 WAAPWLS (default) 0.665
11 puniform (star) 0.553 11 puniform (star) 0.553
12 PEESE (default) 0.461 12 MAN (default) 0.494
13 MAIVE (WAIVE) 0.444 13 PEESE (default) 0.461
14 WILS (default) 0.427 14 MAIVE (WAIVE) 0.444
15 MAN (default) 0.348 15 WILS (default) 0.427
16 AK (AK1) 0.242 16 RTMA (relaxed) 0.308
17 RTMA (relaxed) 0.214 17 AK (AK1) 0.241
18 trimfill (default) 0.193 18 trimfill (default) 0.193
19 WLS (default) 0.154 19 WLS (default) 0.154
20 RMA (default) 0.088 20 RMA (default) 0.088
21 SM (4PSM) 0.057 21 SM (4PSM) 0.072
22 FMA (default) 0.053 22 FMA (default) 0.053
23 mean (default) 0.019 23 mean (default) 0.019
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 WAAPWLS (default) -4.827 1 WAAPWLS (default) -4.827
2 PETPEESE (default) -4.469 2 PETPEESE (default) -4.469
3 EK (default) -4.350 3 EK (default) -4.350
4 PET (default) -4.350 4 PET (default) -4.350
5 PEESE (default) -4.272 5 PEESE (default) -4.272
6 MMPH (default) -3.914 6 AK (AK2) -4.067
7 puniform (default) -3.630 7 MMPH (default) -3.905
8 WLS (default) -2.606 8 puniform (default) -3.629
9 SM (3PSM) -2.571 9 WLS (default) -2.606
10 MAIVE (default) -2.525 10 SM (3PSM) -2.590
11 trimfill (default) -2.517 11 MAIVE (default) -2.525
12 WILS (default) -2.389 12 trimfill (default) -2.519
13 AK (AK2) -2.109 13 WILS (default) -2.389
14 FMA (default) -2.070 14 RTMA (relaxed) -2.079
15 puniform (star) -2.025 15 FMA (default) -2.070
16 RoBMA (PSMA) -1.921 16 puniform (star) -2.025
17 AK (AK1) -1.790 17 RoBMA (PSMA) -1.913
18 RMA (default) -1.448 18 AK (AK1) -1.793
19 mean (default) -0.850 19 MAN (default) -1.617
20 MAN (default) -0.620 20 RMA (default) -1.448
21 RTMA (relaxed) -0.608 21 mean (default) -0.850
22 SM (4PSM) -0.444 22 SM (4PSM) -0.550
23 MAIVE (WAIVE) -0.120 23 MAIVE (WAIVE) -0.120
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 AK (AK2) 0.169 1 MAIVE (WAIVE) 0.226
2 MAIVE (WAIVE) 0.226 2 RoBMA (PSMA) 0.333
3 RoBMA (PSMA) 0.324 3 MAIVE (default) 0.342
4 MAIVE (default) 0.342 4 PETPEESE (default) 0.357
5 PETPEESE (default) 0.357 5 PET (default) 0.397
6 MMPH (default) 0.369 6 EK (default) 0.397
7 PET (default) 0.397 7 puniform (default) 0.483
8 EK (default) 0.397 8 AK (AK2) 0.492
9 puniform (default) 0.483 9 SM (3PSM) 0.493
10 SM (3PSM) 0.491 10 MMPH (default) 0.515
11 MAN (default) 0.503 11 WAAPWLS (default) 0.515
12 WAAPWLS (default) 0.515 12 puniform (star) 0.518
13 puniform (star) 0.518 13 MAN (default) 0.581
14 RTMA (relaxed) 0.538 14 WILS (default) 0.601
15 WILS (default) 0.601 15 SM (4PSM) 0.625
16 SM (4PSM) 0.623 16 PEESE (default) 0.684
17 PEESE (default) 0.684 17 RTMA (relaxed) 0.754
18 trimfill (default) 0.857 18 trimfill (default) 0.857
19 AK (AK1) 0.867 19 AK (AK1) 0.867
20 WLS (default) 0.876 20 WLS (default) 0.876
21 RMA (default) 0.932 21 RMA (default) 0.932
22 FMA (default) 0.954 22 FMA (default) 0.954
23 mean (default) 0.983 23 mean (default) 0.983
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 mean (default) 1.000 1 mean (default) 1.000
2 FMA (default) 0.999 2 FMA (default) 0.999
3 RMA (default) 0.998 3 RMA (default) 0.998
4 trimfill (default) 0.995 4 trimfill (default) 0.995
5 WLS (default) 0.994 5 WLS (default) 0.994
6 AK (AK1) 0.990 6 AK (AK1) 0.990
7 PEESE (default) 0.981 7 PEESE (default) 0.981
8 WAAPWLS (default) 0.948 8 RTMA (relaxed) 0.971
9 puniform (default) 0.923 9 WAAPWLS (default) 0.948
10 MMPH (default) 0.922 10 MMPH (default) 0.928
11 PETPEESE (default) 0.915 11 puniform (default) 0.923
12 EK (default) 0.900 12 PETPEESE (default) 0.915
13 PET (default) 0.900 13 EK (default) 0.900
14 WILS (default) 0.861 14 PET (default) 0.900
15 SM (3PSM) 0.821 15 AK (AK2) 0.879
16 puniform (star) 0.755 16 WILS (default) 0.861
17 AK (AK2) 0.716 17 MAN (default) 0.837
18 MAIVE (default) 0.716 18 SM (3PSM) 0.826
19 MAN (default) 0.705 19 puniform (star) 0.755
20 SM (4PSM) 0.670 20 MAIVE (default) 0.716
21 RTMA (relaxed) 0.634 21 SM (4PSM) 0.684
22 RoBMA (PSMA) 0.629 22 RoBMA (PSMA) 0.634
23 MAIVE (WAIVE) 0.288 23 MAIVE (WAIVE) 0.288
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Subset: High Questionable Research Practices

These results are based on Carter (2019) data-generating mechanism with a total of 252 conditions.

Average Performance

Method performance measures are aggregated across all simulated conditions to provide an overall impression of method performance. However, keep in mind that a method with a high overall ranking is not necessarily the “best” method for a particular application. To select a suitable method for your application, consider also non-aggregated performance measures in conditions most relevant to your application.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MMPH (default) 0.104 1 AK (AK1) 0.105
2 AK (AK1) 0.105 2 AK (AK2) 0.117
3 PEESE (default) 0.127 3 PEESE (default) 0.127
4 PETPEESE (default) 0.129 4 PETPEESE (default) 0.129
5 pcurve (default) 0.129 5 pcurve (default) 0.129
6 WILS (default) 0.130 6 WILS (default) 0.130
7 WAAPWLS (default) 0.131 7 WAAPWLS (default) 0.131
8 EK (default) 0.140 8 EK (default) 0.140
9 PET (default) 0.140 9 PET (default) 0.140
10 puniform (default) 0.149 10 MMPH (default) 0.141
11 AK (AK2) 0.158 11 puniform (default) 0.149
12 trimfill (default) 0.159 12 trimfill (default) 0.159
13 RoBMA (PSMA) 0.162 13 RoBMA (PSMA) 0.164
14 MAIVE (default) 0.169 14 MAIVE (default) 0.169
15 FMA (default) 0.175 15 FMA (default) 0.175
15 WLS (default) 0.175 15 WLS (default) 0.175
17 MAN (default) 0.215 17 MAN (default) 0.211
18 RMA (default) 0.244 18 RMA (default) 0.244
19 mean (default) 0.358 19 RTMA (relaxed) 0.295
20 MAIVE (WAIVE) 0.378 20 mean (default) 0.358
21 SM (3PSM) 0.493 21 MAIVE (WAIVE) 0.378
22 RTMA (relaxed) 0.495 22 SM (3PSM) 0.468
23 SM (4PSM) 0.736 23 SM (4PSM) 0.779
24 puniform (star) 314.821 24 puniform (star) 314.821

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 WILS (default) -0.012 1 WILS (default) -0.012
2 MAIVE (default) -0.012 2 MAIVE (default) -0.012
3 RTMA (relaxed) 0.013 3 EK (default) -0.032
4 EK (default) -0.032 4 PET (default) -0.032
5 PET (default) -0.032 5 pcurve (default) 0.038
6 pcurve (default) 0.038 6 MAN (default) 0.040
7 MMPH (default) 0.043 7 PETPEESE (default) 0.046
8 PETPEESE (default) 0.046 8 AK (AK2) 0.046
9 AK (AK2) 0.050 9 puniform (default) 0.054
10 puniform (default) 0.054 10 AK (AK1) 0.062
11 AK (AK1) 0.062 11 PEESE (default) 0.075
12 PEESE (default) 0.075 12 MMPH (default) 0.086
13 WAAPWLS (default) 0.101 13 RoBMA (PSMA) -0.094
14 RoBMA (PSMA) -0.103 14 WAAPWLS (default) 0.101
15 trimfill (default) 0.136 15 SM (3PSM) -0.130
16 SM (3PSM) -0.148 16 trimfill (default) 0.136
17 MAN (default) -0.151 17 FMA (default) 0.159
18 FMA (default) 0.159 17 WLS (default) 0.159
18 WLS (default) 0.159 19 RMA (default) 0.237
20 RMA (default) 0.237 20 RTMA (relaxed) 0.242
21 MAIVE (WAIVE) -0.250 21 MAIVE (WAIVE) -0.250
22 SM (4PSM) -0.291 22 SM (4PSM) -0.299
23 mean (default) 0.353 23 mean (default) 0.353
24 puniform (star) -67.517 24 puniform (star) -67.517

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 RMA (default) 0.038 1 RMA (default) 0.038
2 AK (AK1) 0.041 2 AK (AK1) 0.042
3 trimfill (default) 0.042 3 trimfill (default) 0.042
4 mean (default) 0.048 4 mean (default) 0.048
5 pcurve (default) 0.055 5 pcurve (default) 0.055
6 WLS (default) 0.057 6 WLS (default) 0.057
7 FMA (default) 0.057 7 FMA (default) 0.057
8 MMPH (default) 0.059 8 WAAPWLS (default) 0.067
9 WAAPWLS (default) 0.067 9 PEESE (default) 0.067
10 PEESE (default) 0.067 10 puniform (default) 0.070
11 puniform (default) 0.070 11 MMPH (default) 0.074
12 AK (AK2) 0.079 12 AK (AK2) 0.074
13 MAN (default) 0.084 13 PETPEESE (default) 0.091
14 RoBMA (PSMA) 0.090 14 WILS (default) 0.097
15 PETPEESE (default) 0.091 15 PET (default) 0.105
16 WILS (default) 0.097 16 EK (default) 0.105
17 PET (default) 0.105 17 MAN (default) 0.109
18 EK (default) 0.105 18 RoBMA (PSMA) 0.109
19 MAIVE (default) 0.121 19 MAIVE (default) 0.121
20 MAIVE (WAIVE) 0.215 20 RTMA (relaxed) 0.149
21 RTMA (relaxed) 0.303 21 MAIVE (WAIVE) 0.215
22 SM (3PSM) 0.436 22 SM (3PSM) 0.414
23 SM (4PSM) 0.618 23 SM (4PSM) 0.661
24 puniform (star) 305.910 24 puniform (star) 305.910

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MMPH (default) 1.302 1 AK (AK2) 1.346
2 puniform (star) 1.717 2 puniform (star) 1.717
3 AK (AK1) 1.951 3 AK (AK1) 1.953
4 EK (default) 2.025 4 EK (default) 2.025
5 RoBMA (PSMA) 2.053 5 WILS (default) 2.094
6 WILS (default) 2.094 6 PET (default) 2.098
7 PET (default) 2.098 7 RoBMA (PSMA) 2.140
8 SM (3PSM) 2.262 8 SM (3PSM) 2.238
9 PETPEESE (default) 2.446 9 PETPEESE (default) 2.446
10 WAAPWLS (default) 2.507 10 WAAPWLS (default) 2.507
11 puniform (default) 2.546 11 puniform (default) 2.546
12 PEESE (default) 2.736 12 MMPH (default) 2.590
13 MAIVE (default) 2.759 13 PEESE (default) 2.736
14 MAN (default) 2.861 14 MAIVE (default) 2.759
15 AK (AK2) 2.956 15 SM (4PSM) 3.770
16 SM (4PSM) 3.798 16 trimfill (default) 3.996
17 trimfill (default) 3.991 17 MAN (default) 4.399
18 WLS (default) 4.658 18 WLS (default) 4.658
19 FMA (default) 5.182 19 FMA (default) 5.182
20 MAIVE (WAIVE) 5.601 20 MAIVE (WAIVE) 5.601
21 RMA (default) 7.051 21 RMA (default) 7.051
22 RTMA (relaxed) 9.342 22 RTMA (relaxed) 7.172
23 mean (default) 11.423 23 mean (default) 11.423
24 pcurve (default) NaN 24 pcurve (default) NaN

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 SM (3PSM) 0.672 1 AK (AK2) 0.691
2 SM (4PSM) 0.663 2 SM (3PSM) 0.659
3 MMPH (default) 0.654 3 SM (4PSM) 0.654
4 RoBMA (PSMA) 0.648 4 RoBMA (PSMA) 0.645
5 RTMA (relaxed) 0.645 5 puniform (star) 0.641
6 puniform (star) 0.641 6 MMPH (default) 0.525
7 AK (AK2) 0.596 7 puniform (default) 0.510
8 puniform (default) 0.510 8 AK (AK1) 0.504
9 MAN (default) 0.510 9 MAIVE (default) 0.474
10 AK (AK1) 0.505 10 WAAPWLS (default) 0.444
11 MAIVE (default) 0.474 11 WILS (default) 0.432
12 WAAPWLS (default) 0.444 12 PETPEESE (default) 0.409
13 WILS (default) 0.432 13 EK (default) 0.396
14 PETPEESE (default) 0.409 14 PET (default) 0.382
15 EK (default) 0.396 15 MAIVE (WAIVE) 0.370
16 PET (default) 0.382 16 trimfill (default) 0.360
17 MAIVE (WAIVE) 0.370 17 PEESE (default) 0.329
18 trimfill (default) 0.361 18 MAN (default) 0.288
19 PEESE (default) 0.329 19 WLS (default) 0.218
20 WLS (default) 0.218 20 RTMA (relaxed) 0.191
21 FMA (default) 0.179 21 FMA (default) 0.179
22 RMA (default) 0.149 22 RMA (default) 0.149
23 mean (default) 0.040 23 mean (default) 0.040
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 FMA (default) 0.088 1 FMA (default) 0.088
2 WLS (default) 0.122 2 WLS (default) 0.122
3 PEESE (default) 0.126 3 PEESE (default) 0.126
4 trimfill (default) 0.144 4 trimfill (default) 0.144
5 mean (default) 0.147 5 mean (default) 0.147
6 RMA (default) 0.147 6 RMA (default) 0.147
7 PETPEESE (default) 0.148 7 PETPEESE (default) 0.148
8 AK (AK1) 0.155 8 AK (AK1) 0.155
9 WAAPWLS (default) 0.163 9 WAAPWLS (default) 0.163
10 WILS (default) 0.186 10 WILS (default) 0.186
11 PET (default) 0.197 11 PET (default) 0.197
12 EK (default) 0.213 12 EK (default) 0.213
13 MMPH (default) 0.248 13 MAN (default) 0.233
14 puniform (default) 0.270 14 MMPH (default) 0.236
15 MAIVE (default) 0.296 15 AK (AK2) 0.252
16 RoBMA (PSMA) 0.337 16 puniform (default) 0.270
17 MAN (default) 0.417 17 MAIVE (default) 0.296
18 puniform (star) 0.422 18 RoBMA (PSMA) 0.333
19 MAIVE (WAIVE) 0.524 19 puniform (star) 0.422
20 AK (AK2) 0.725 20 MAIVE (WAIVE) 0.524
21 SM (3PSM) 1.035 21 RTMA (relaxed) 0.568
22 SM (4PSM) 1.711 22 SM (3PSM) 0.907
23 RTMA (relaxed) 5.805 23 SM (4PSM) 1.606
24 pcurve (default) NaN 24 pcurve (default) NaN

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 puniform (default) 1.992 1 puniform (default) 1.992
2 RoBMA (PSMA) 1.581 2 RoBMA (PSMA) 1.344
3 MMPH (default) 1.174 3 MAIVE (default) 0.880
4 MAIVE (default) 0.880 4 PET (default) 0.804
5 PET (default) 0.804 5 EK (default) 0.802
6 EK (default) 0.802 6 PETPEESE (default) 0.739
7 SM (3PSM) 0.781 7 SM (3PSM) 0.737
8 PETPEESE (default) 0.739 8 MMPH (default) 0.581
9 AK (AK2) 0.451 9 MAN (default) 0.535
10 WILS (default) 0.435 10 WILS (default) 0.435
11 puniform (star) 0.400 11 puniform (star) 0.400
12 MAN (default) 0.393 12 AK (AK2) 0.241
13 WAAPWLS (default) 0.236 13 WAAPWLS (default) 0.236
14 SM (4PSM) 0.153 14 SM (4PSM) 0.157
15 PEESE (default) 0.093 15 PEESE (default) 0.093
16 AK (AK1) 0.071 16 AK (AK1) 0.071
17 trimfill (default) 0.035 17 RTMA (relaxed) 0.067
18 WLS (default) 0.034 18 trimfill (default) 0.035
19 RMA (default) 0.013 19 WLS (default) 0.034
20 FMA (default) 0.007 20 RMA (default) 0.013
21 mean (default) 0.000 21 FMA (default) 0.007
22 MAIVE (WAIVE) -0.094 22 mean (default) 0.000
23 RTMA (relaxed) -0.192 23 MAIVE (WAIVE) -0.094
24 pcurve (default) NaN 24 pcurve (default) NaN

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Log Value Rank Method Log Value
1 PETPEESE (default) -4.432 1 PETPEESE (default) -4.432
2 PET (default) -4.420 2 PET (default) -4.420
3 EK (default) -4.420 3 EK (default) -4.420
4 WAAPWLS (default) -3.509 4 WAAPWLS (default) -3.509
5 puniform (default) -3.426 5 puniform (default) -3.426
6 MMPH (default) -3.322 6 MMPH (default) -3.137
7 SM (3PSM) -2.241 7 AK (AK2) -3.133
8 MAIVE (default) -2.230 8 MAN (default) -2.300
9 PEESE (default) -2.108 9 SM (3PSM) -2.265
10 AK (AK2) -1.792 10 MAIVE (default) -2.230
11 RoBMA (PSMA) -1.576 11 PEESE (default) -2.108
12 puniform (star) -1.566 12 puniform (star) -1.566
13 WLS (default) -1.374 13 RoBMA (PSMA) -1.560
14 WILS (default) -1.274 14 WLS (default) -1.374
15 trimfill (default) -0.981 15 WILS (default) -1.274
16 AK (AK1) -0.972 16 RTMA (relaxed) -1.008
17 FMA (default) -0.821 17 trimfill (default) -0.982
18 SM (4PSM) -0.775 18 AK (AK1) -0.973
19 RMA (default) -0.613 19 SM (4PSM) -0.872
20 MAN (default) -0.593 20 FMA (default) -0.821
21 mean (default) -0.213 21 RMA (default) -0.613
22 MAIVE (WAIVE) 0.267 22 mean (default) -0.213
23 RTMA (relaxed) 0.677 23 MAIVE (WAIVE) 0.267
24 pcurve (default) NaN 24 pcurve (default) NaN

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 MAIVE (WAIVE) 0.361 1 MAIVE (WAIVE) 0.361
2 RoBMA (PSMA) 0.369 2 RoBMA (PSMA) 0.374
3 MAIVE (default) 0.399 3 MAIVE (default) 0.399
4 MMPH (default) 0.474 4 puniform (star) 0.492
5 SM (3PSM) 0.486 5 SM (3PSM) 0.497
6 puniform (star) 0.492 6 puniform (default) 0.511
7 puniform (default) 0.511 7 PET (default) 0.518
8 PET (default) 0.518 8 EK (default) 0.518
9 EK (default) 0.518 9 PETPEESE (default) 0.549
10 MAN (default) 0.533 10 WILS (default) 0.569
11 PETPEESE (default) 0.549 11 SM (4PSM) 0.599
12 WILS (default) 0.569 12 MMPH (default) 0.640
13 SM (4PSM) 0.593 13 MAN (default) 0.695
14 RTMA (relaxed) 0.603 14 WAAPWLS (default) 0.783
15 AK (AK2) 0.637 15 AK (AK2) 0.860
16 WAAPWLS (default) 0.783 16 PEESE (default) 0.916
17 PEESE (default) 0.916 17 RTMA (relaxed) 0.925
18 AK (AK1) 0.945 18 AK (AK1) 0.945
19 trimfill (default) 0.968 19 WLS (default) 0.968
20 WLS (default) 0.968 20 trimfill (default) 0.968
21 RMA (default) 0.988 21 RMA (default) 0.988
22 FMA (default) 0.992 22 FMA (default) 0.992
23 mean (default) 0.999 23 mean (default) 0.999
24 pcurve (default) NaN 24 pcurve (default) NaN

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Conditional on Convergence
Replacement if Non-Convergence
Rank Method Value Rank Method Value
1 mean (default) 1.000 1 mean (default) 1.000
2 RMA (default) 1.000 2 RMA (default) 1.000
3 FMA (default) 1.000 3 FMA (default) 1.000
4 trimfill (default) 0.998 4 trimfill (default) 0.998
5 WLS (default) 0.998 5 WLS (default) 0.998
6 AK (AK1) 0.998 6 AK (AK1) 0.998
7 PEESE (default) 0.994 7 AK (AK2) 0.994
8 WAAPWLS (default) 0.969 8 PEESE (default) 0.994
9 AK (AK2) 0.959 9 RTMA (relaxed) 0.984
10 PETPEESE (default) 0.939 10 WAAPWLS (default) 0.969
11 puniform (default) 0.930 11 PETPEESE (default) 0.939
12 EK (default) 0.925 12 MMPH (default) 0.932
13 PET (default) 0.925 13 puniform (default) 0.930
14 MMPH (default) 0.921 14 EK (default) 0.925
15 WILS (default) 0.801 15 PET (default) 0.925
16 SM (3PSM) 0.777 16 MAN (default) 0.889
17 MAIVE (default) 0.754 17 WILS (default) 0.801
18 MAN (default) 0.711 18 SM (3PSM) 0.789
19 SM (4PSM) 0.689 19 MAIVE (default) 0.754
20 puniform (star) 0.671 20 SM (4PSM) 0.704
21 RoBMA (PSMA) 0.624 21 puniform (star) 0.671
22 RTMA (relaxed) 0.509 22 RoBMA (PSMA) 0.629
23 MAIVE (WAIVE) 0.270 23 MAIVE (WAIVE) 0.270
24 pcurve (default) NaN 24 pcurve (default) NaN

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Conditional on Method Convergence)

The results below are conditional on method convergence. Note that the methods might differ in convergence rate and are therefore not compared on the same data sets.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

By-Condition Performance (Replacement in Case of Non-Convergence)

The results below incorporate method replacement to handle non-convergence. If a method fails to converge, its results are replaced with the results from a simpler method (e.g., random-effects meta-analysis without publication bias adjustment). This emulates what a data analyst may do in practice in case a method does not converge. However, note that these results do not correspond to “pure” method performance as they might combine multiple different methods. See Method Replacement Strategy for details of the method replacement specification.

Raincloud plot showing convergence rates across different methods

Raincloud plot showing RMSE (Root Mean Square Error) across different methods

RMSE (Root Mean Square Error) is an overall summary measure of estimation performance that combines bias and empirical SE. RMSE is the square root of the average squared difference between the meta-analytic estimate and the true effect across simulation runs. A lower RMSE indicates a better method. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing bias across different methods

Bias is the average difference between the meta-analytic estimate and the true effect across simulation runs. Ideally, this value should be close to 0. Values lower than -0.5 or larger than 0.5 are visualized as -0.5 and 0.5 respectively.

Raincloud plot showing bias across different methods

The empirical SE is the standard deviation of the meta-analytic estimate across simulation runs. A lower empirical SE indicates less variability and better method performance. Values larger than 0.5 are visualized as 0.5.

Raincloud plot showing 95% confidence interval width across different methods

The interval score measures the accuracy of a confidence interval by combining its width and coverage. It penalizes intervals that are too wide or that fail to include the true value. A lower interval score indicates a better method. Values larger than 100 are visualized as 100.

Raincloud plot showing 95% confidence interval coverage across different methods

95% CI coverage is the proportion of simulation runs in which the 95% confidence interval contained the true effect. Ideally, this value should be close to the nominal level of 95%.

Raincloud plot showing 95% confidence interval width across different methods

95% CI width is the average length of the 95% confidence interval for the true effect. A lower average 95% CI length indicates a better method.

Raincloud plot showing positive likelihood ratio across different methods

The positive likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a positive likelihood ratio greater than 1 (or a log positive likelihood ratio greater than 0). A higher (log) positive likelihood ratio indicates a better method.

Raincloud plot showing negative likelihood ratio across different methods

The negative likelihood ratio is an overall summary measure of hypothesis testing performance that combines power and type I error rate. It indicates how much a non-significant test result changes the odds of the alternative hypothesis versus the null hypothesis. A useful method has a negative likelihood ratio less than 1 (or a log negative likelihood ratio less than 0). A lower (log) negative likelihood ratio indicates a better method.

Raincloud plot showing Type I Error rates across different methods

The type I error rate is the proportion of simulation runs in which the null hypothesis of no effect was incorrectly rejected when it was true. Ideally, this value should be close to the nominal level of 5%.

Raincloud plot showing statistical power across different methods

The power is the proportion of simulation runs in which the null hypothesis of no effect was correctly rejected when the alternative hypothesis was true. A higher power indicates a better method.

Session Info

This report was compiled on Wed Sep 30 15:45:59 2026 (UTC) using the following computational environment

## R version 4.6.1 (2026-06-24)
## Platform: x86_64-pc-linux-gnu
## Running under: Ubuntu 24.04.5 LTS
## 
## Matrix products: default
## BLAS:   /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3 
## LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.26.so;  LAPACK version 3.12.0
## 
## locale:
##  [1] LC_CTYPE=C.UTF-8       LC_NUMERIC=C           LC_TIME=C.UTF-8       
##  [4] LC_COLLATE=C.UTF-8     LC_MONETARY=C.UTF-8    LC_MESSAGES=C.UTF-8   
##  [7] LC_PAPER=C.UTF-8       LC_NAME=C              LC_ADDRESS=C          
## [10] LC_TELEPHONE=C         LC_MEASUREMENT=C.UTF-8 LC_IDENTIFICATION=C   
## 
## time zone: UTC
## tzcode source: system (glibc)
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
## [1] scales_1.4.0                   ggdist_3.3.3                  
## [3] ggplot2_4.0.3                  PublicationBiasBenchmark_0.3.0
## 
## loaded via a namespace (and not attached):
##  [1] gtable_0.3.6         xfun_0.61            bslib_0.12.0        
##  [4] htmlwidgets_1.6.4    lattice_0.22-9       vctrs_0.7.3         
##  [7] tools_4.6.1          Rdpack_2.6.6         generics_0.1.4      
## [10] curl_8.0.0           sandwich_3.1-3       tibble_3.3.1        
## [13] pkgconfig_2.0.3      RColorBrewer_1.1-3   S7_0.2.2            
## [16] desc_1.4.3           distributional_0.9.0 lifecycle_1.0.5     
## [19] compiler_4.6.1       farver_2.1.2         stringr_1.6.0       
## [22] textshaping_1.0.5    htmltools_0.5.9      sass_0.4.10         
## [25] clubSandwich_0.7.0   yaml_2.3.12          pillar_1.11.1       
## [28] pkgdown_2.2.1        jquerylib_0.1.4      cachem_1.1.0        
## [31] tidyselect_1.2.1     digest_0.6.39        stringi_1.8.9       
## [34] dplyr_1.2.1          purrr_1.2.2          labeling_0.4.3      
## [37] fastmap_1.2.0        grid_4.6.1           cli_3.6.6           
## [40] magrittr_2.0.5       triebeard_0.4.1      crul_1.6.0          
## [43] osfr_0.2.9           withr_3.0.3          rmarkdown_2.32      
## [46] httr_1.4.9           otel_0.2.0           ragg_1.5.2          
## [49] zoo_1.9-1            kableExtra_1.4.1     memoise_2.0.1       
## [52] evaluate_1.0.5       knitr_1.52           rbibutils_2.4.1     
## [55] viridisLite_0.4.3    rlang_1.3.0          urltools_1.7.3.1    
## [58] Rcpp_1.1.2           glue_1.8.1           httpcode_0.3.0      
## [61] xml2_1.6.0           svglite_2.2.2        rstudioapi_0.19.0   
## [64] jsonlite_2.0.0       R6_2.6.1             systemfonts_1.3.2   
## [67] fs_2.1.0