本包的基本核心仍然是mixed.solve,它求解了除残差之外的一个方差分量的混合模型。该函数估计两个方差分量,通过ML或REML模型估计,估计的执行使用Kang等(2008)描述的光谱分析算法。在这个过程中,很容易创建表型协方差矩阵的逆矩阵,逆矩阵随后用于固定效应和随机效应的BLUE和BLUP求解计算(Searle et al. 1992)。在(Endelman 2011)中,作者展示了mixed.solve如何用于全基因组预测,要么建模标记作为随机效应,要么家系随机效应。
用均值补缺失肯定会更快,在很多情况下,在GEBV的预测精度上,均值的表现与EM算法和其他更先进的方法一样好。然而,与EM方法相比,均值的方法的育种值往往更有偏向性(Poland et al. 2012)。对两个矩阵的平均对角线元素比较表明,EM算法的结果更加接近给定1%杂合率的期望,表达式1+f≈2。提出
round(mean(diag(A2)),2) # imputed with mean
round(mean(dia(A1)),2) # imputed with EM
A矩阵的收缩估计
A.mat的另一个特性是收缩估计,它可以用于低密度的标记,例如来自384个SNP的芯片。当家系的数量与标记的数量相当或更多时,上面的方程可能是最优化的亲缘关系矩阵估计。Endelman and Jannink (2012)建议将估算值降低到(1+f)I,用收缩强度选择最小化均方误差。
test <- which(sets==1)
yNA <- Y[,1] # grain yield in environment 1
yNA[test] <- NA # mask yields for validation set
data1 <- data.frame(y=yNA, gid=1:599)
Endelman, J.B. 2011. Ridge regression and other kernels for genomic selection with R package rrBLUP. Plant Genome 4:250–255. doi:10.3835/plantgenome2011.08.0024
Endelman, J.B., and J.-L. Jannink. 2012. Shrinkage estimation of the realized relationship matrix. G3:Genes, Genomes, Genetics. 2:1405-1413. doi:10.1534/g3.112.004259
Kang et al. 2008. Efficient control of population structure in model organism association mapping. Genetics 178:1709–1723.
Pérez et al. 2010. Genomic-enabled prediction based on molecular markers and pedigree using the Bayesian Linear Regression package in R. Plant Genome 3:106–116.
Poland, J., J. Endelman et al. 2012. Genomic selection in wheat breeding using genotyping-by-sequencing. Plant Genome 5:103–113. doi: 10.3835/plantgenome2012.06.0006.
Searle et al. 1992. Variance Components. John Wiley & Sons, Hoboken.
VanRaden, P.M. 2008. Efficient methods to compute genomic predictions. J. Dairy Science 91:4414–4423.