Showing posts with label writings. Show all posts
Showing posts with label writings. Show all posts

2015-08-13

Reliability of pedigree-based and genomic evaluations in selected populations

Paper on "Reliability of pedigree-based and genomic evaluations in selected populations" has finally been published after tedious reviews. This work is a follow up study on reliability of genetic evaluation in selected populations by Piter Bijma (link), but this time connecting pedigree-based and genomic evaluations. The take-home message is that PEV-based reliabilities of genomic predictions are not so much affected by selection as pedigree predictions. This has implications when different breeding program designs are compared using PEV reliabilities - commonly PEV reliability of genomic and pedigree predictions are compared - and suggests that genotyping female selection candidates has been undervalued and should be reconsidered.

2015-03-07

Potential of genotyping-by-sequencing for genomic selection in livestock populations

Our new paper titled "Potential of genotyping-by-sequencing for genomic selection in livestock populations" has been published in Genetics Selection Evolution. This work shows that genotypes called from low-coverage sequencing data can be equally or even more powerful for genomic prediction than high-quality SNP genotypes. In particular manipulation of coverage allows us to increase number of genotyped individuals at the expense of genotype quality and this can bring us a long way before accuracy of genomic predictions falls significantly. Another useful application is in increasing selection intensity by genotyping more/all selection candidates.

2013-01-27

The contribution of dominance and inbreeding depression in estimating variance components for litter size in Pannon White rabbits

Our paper on analysis of dominance in Pannon White rabbits has been accepted and will appear in Journal of Animal breeding and Genetics. Abstract says:

In a synthetic closed population of Pannon White rabbits, additive (VA), dominance (VD) and permanent environmental (VPe) variance components as well as doe (bFd) and litter (bFl) inbreeding depression were estimated for the number of kits born alive (NBA), number of kits born dead (NBD) and total number of kits born (TNB). The data set consisted of 18,398 kindling records of 3883 does collected from 1992 to 2009. Six models were used to estimate dominance and inbreeding effects. The most complete model estimated VA and VD to contribute 5.5 ± 1.1% and 4.8 ± 2.4%, respectively, to total phenotypic variance (VP) for NBA; the corresponding values for NBD were 1.9 ± 0.6% and 5.3 ± 2.4%, for TNB, 6.2 ± 1.0% and 8.1 ± 3.2% respectively. These results indicate the presence of considerable VD. Including dominance in the model generally reduced VA and VPe estimates, and had only a very small effect on inbreeding depression estimates. Including inbreeding covariates did not affect estimates of any variance component. A 10% increase in doe inbreeding significantly increased NBD (bFd = 0.18 ± 0.07), while a 10% increase in litter inbreeding significantly reduced NBA (bFl = −0.41 ± 0.11) and TNB (bFl = −0.34 ± 0.10). These findings argue for including dominance effects in models of litter size traits in populations that exhibit significant dominance relationships.

The contribution of dominance and inbreeding depression in estimating variance components for litter size i... by

Genomic evaluations using similarity between haplotypes

Our paper on using haplotypes to build relationships between individuals has been accepted and will appear in Journal of Animal breeding and Genetics. Abstract says:

Long-range phasing and haplotype library imputation methodologies are accurate and efficient methods to provide haplotype information that could be used in prediction of breeding value or phenotype. Modelling long haplotypes as independent effects in genomic prediction would be inefficient due to the many effects that need to be estimated and phasing errors, even if relatively low in frequency, exacerbate this problem. One approach to overcome this is to use similarity between haplotypes to model covariance of genomic effects by region or of animal breeding values. We developed a simple method to do this and tested impact on genomic prediction by simulation. Results show that the diagonal and off-diagonal elements of a genomic relationship matrix constructed using the haplotype similarity method had higher correlations with the true relationship between pairs of individuals than genomic relationship matrices built using unphased genotypes or assumed unrelated haplotypes. However, the prediction accuracy of such haplotype-based prediction methods was not higher than those based on unphased genotype information.

Genomic evaluations using similarity between haplotypes by

2013-01-25

Simulation of different strategies to implement genomic selection in plant breeding

Our submitted abstract for the VIPCA conference.

Gorjanc G, Crossa J, Dreisigacker S, Autrique E, Hickey J

Genomic information provides a rich resource for plant improvement. The greatest advantage of genomic information is achived when used early in a programme. The aim of this contribution was to present different ways of utilizing genomic information in cross and self pollinated species early in the bi-parental population (BP) cycle in the F2 generation by simulation. Simulation mimicked a real breeding programme with BP of different degree of relationship. Simulated trait heritability was 0.5. Predictions were always within a focal BP (BPX) avoiding the effect of inflated accuracy due to population structure. Measure of interest was the correlation between true and predicted additive genetic value (accuracy). Predictions were based on training on collected data within BPX, from BP having one parent in common with BPX (BPPC), from BP having one grand-parent in common with BPX (BPGC), and from unrelated BP (BPUR). Results show that genomic predictions can be accurate early in the cycle. Accurate predictions within the BPX can be obtained with a small number of markers and moderate phenotyping of F2 derived lines. With 200 markers and 50 phenotypes accuracy was around 0.60 and increased to 0.80 with 100 phenotypes. Use of information from relatives requires denser marker panels and more phenotypes to properly model the expanding number of linkage blocks in progressively distantly related BP. However, these requirements can be spread over many related BP reducing the cost per BP. For example accuracy of 0.6 can be achieved by using 1,000 markers and 50 phenotypes from 4 BPPC and 40 BPUR or by using 10,000 markers and 50 phenotypes from 8 BPGC and 50 BPUR. Requirement for denser marker panels increases needed investment, which can be offset by proper use of imputation method making the whole procedure cost effective for real applications.

2012-04-17

Simulated data for genomic selection and GWAS using a combination of coalescent and gene drop methods

Our (Hickey and Gorjanc) paper describing simulation using coalescent and gene drop methods has been published in the new open G3 journal. This contribution is a part of the Genomic selection collection of the Genetic Society of America that is nicely described here

2011-10-05

Ovčereja na Norveškem

Ovčereja na Norveškem

2011-10-04

ASD 2011 - Primošten

Animal Science Days (ASD) conference was held in Primošten this year. I participated with the following contributions as a co-author:
  • Lambing Interval in Jezersko-Solčava and Improved Jezersko-Solčava Breed. PDF
  • Partitioning of Genetic Trends by Originin Croatian Simmental Cattle. PDF
  • Estimation of Variance Components for Litter Size in the First and Later Parities in Improved Jezersko-Solcava Sheep. PDF
  • The Effect of Coenzyme Q10 and Lipoic Acid Added to the Feed of Hens on Physical Characteristics of Eggs. PDF

2011-09-30

InterBull: Partitioning of international genetic trends by origin in Brown Swiss bulls

I attended the InterBull meeting this year in Stavanger (Norway), which was jointly organised with the EAAP conference in the same place. I participated with contribution titled "Partitioning of international genetic trends by origin in Brown Swiss bulls" co-authored with colleagues. In essence we partitioned  breeding values of Brown Swiss animals by origin of selection and summarized those partitions as origin specific genetic trends. This gives us an opportunity to evaluate the effect of selection performed in different countries and how this affects the global genetic trends. Results for this breed are quite shocking! Look into the paper and talk (see bellow) for more ;)

We were looking forward for comments and/or critiques about the applied method from the audience, but did not get any direct questions. Several people approached me after the talk and one of the comments was that our method is nothing more than multiplying breeding values by the share of genes coming from different origin. This is not true and I will demonstrate this in with a simple example.

Let us assume that we have a simple small pedigree as shown bellow with R code, where column names are: id =individual code, fid = father code, mid = mother code, ori = origin, and bv = breeding value. This pedigree could represent situation where we constantly use foreign sires - here only two generations are being shown. Origin represents the country of registering/selecting aninmal.

## Simple example
example <- data.frame( id=c(  1,   2,   3,   4,   5),
                      fid=c( NA,  NA,   1,  NA,   3),
                      mid=c( NA,  NA,   2,  NA,   4),
                      ori=c("A", "B", "A", "B", "A"),
                       bv=c(100, 106, 104, 106, 103))

Now we would like to perform gene proportion analysis and partitioning of supplied breeding values, both according to origin. This is easy to achieve "by hand" for this simple example (see paper bellow for maths), but tedious for bigger examples. I have wrote an R package partAGV (not yet publicly available, but you can contact me) that can be used for such analyses.

## For gene proportion analysis
example$gp <- 1
 
library(partAGV)
partAGV(example, colAGV=6:5)
 
## Gene proportions
##    id  fid  mid ori gp gp_pa gp_w gp_A gp_B
## 1   1 <NA> <NA>   A  1     0    1 1.00 0.00
## 2   2 <NA> <NA>   B  1     0    1 0.00 1.00
## 3   3    1    2   A  1     1    0 0.50 0.50
## 4   4 <NA> <NA>   B  1     0    1 0.00 1.00
## 5   5    3    4   A  1     1    0 0.25 0.75
 
## Partitions of breeding values 
##    id  fid  mid ori  bv bv_pa bv_w  bv_A  bv_B
## 1   1 <NA> <NA>   A 100     0  100 100.0   0.0
## 2   2 <NA> <NA>   B 106     0  106   0.0 106.0
## 3   3    1    2   A 104   103    1  51.0  53.0
## 4   4 <NA> <NA>   B 106     0  106   0.0 106.0
## 5   5    3    4   A 103   105   -2  23.5  79.5

As we can see gene proportions are as expected: 1/2 for each origin in individual 3 and 1/4 vs. 3/4 for individual 5. Partitions of breeding values show that in animal 5 we 23.5 (out of 105) is attributed to selection work done in country A and 79.5 is attributed to selection work done in country B. Now, if we multiply breeding value of individual 5 (105) by origin specific gene proportions (1/4 and 3/4) we get 25.75 and 77.5, which is similar to partitions, but not the same. This shows that our method enables separation of gene flow and selection work preformed by particular country. In order to see this algebraically we can write breeding values of this individual as:

a_5   = 1/2 a_3 + 1/2 a_4 + w_5
      = 1/2 (1/2 a_1 + 1/2 a_2   +     w_3)  + 1/2 w_4 + w_5
      =      1/4 a_1 + 1/4 a_2   + 1/2 w_3   + 1/2 w_4 + w_5
      =      1/4 w_1 + 1/4 w_2   + 1/2 w_3   + 1/2 w_4 + w_5
      =      1/4 100 + 1/4 106   + 1/2   1   + 1/2 106 +  -2
      =           25 +      26.5 +       0.5 +      53 +  -2

If we now collect terms specific to each origin we get:


a_5_A =           25 +                   0.5 +         +  -2 = 23.5
a_5_B =                     26.5 +                  53       = 79.5



Partitioning of international genetic trends by origin in Brown Swiss bulls

2010-12-31

Simple reparameterization to improve convergence in linear mixed models

Quite some time ago I had an opportunity to visit Luis Alberto Garcia-Cortes at INIA (Madrid, Spain), where I was introduced to Bayesian statistics and McMC methods. We were also doing some research in that area. We started to write a short note, which has been rejected for publication due to lack of "breadth". We failed to find time to work on this topic further on. Therefore, I published the note in our local journal. The note can be found here or embeded bellow.
Simple reparameterization to improve convergence in linear mixed models

2010-02-27

9th WCGALP: Flexible Bayesian Inference of Animal Model Parameters Using BUGS Program

Conference madness is continuing ;) Bellow is my contribution for 9th WCGALP (World Congress on Genetics Applied to Livestock Production), which is held every four years. The above site describes it as: "This congress is the premier meeting point for scientists around the world involved in genetic improvement of livestock". My contribution is again on fitting so called animal model (pedigree based mixed model) in BUGS. The contributions must be very short (only four pages), so there is not much to show. I hope the contribution is going to be accepted so that I can spread this idea among the animal breeders.
Flexible Bayesian Inference of Animal Model Parameters Using BUGS Program

2009-07-03

Evaluation of different approaches for the estimation of daily yield from single milk testing scheme in cattle

Janez finnished and submitted the paper "Evaluation of different approaches for the estimation of daily yield from single milk testing scheme in cattle". He did the majority of job! I improved his work a bit with the comments and restructuring/rewritting some parts of the paper.


Evaluation of different approaches for the estimation of daily yield from single milk testing scheme in cattle

2009-05-25

Fitting pedigree based mixed models in BUGS software

Bellow is my abstract for talk at joint congress of SBD and SGD at Otočec. I will show how BUGS (Bayesian Using Gibbs Sampling) software can be used to fit the so called animal model. Now I need to prepair the talk and perhaps even write a short communication for some journal. Added 2009-09-20: The talk is available here.
Pedigree based mixed model (commonly called animal model) is an important class of statistical models for inference of quantitative genetic parameters in various fields such as animal and plant breeding, evolutionary biology and human genetics. In last years Bayesian statistics has been introduced to a set of standard statistical procedures of a quantitative geneticists' toolbox, due to the ever increasing complexity of fitted models. While several specific programs can be used to fit animal model using Bayesian approach, none of them provide and easy to use and flexible environment for the development and testing of models. It is common to use favourite programming language to develop the needed programs, but this requires a considerable amount of programming and statistical skills. A viable alternative is to use general purpose statistical packages. BUGS (Bayesian Using Gibbs Sampling) is a popular and fairly flexible program for the Bayesian analysis of complex statistical models using Markov Chain Monte Carlo methods. Recently two reports of fitting animal model in BUGS were given, but both failed to provide a generic procedure that can be used independently of the collected data. Here, a generic description of animal model is presented using the concept of graphical models. This description was translated to BUGS language and fitted to a small example. Comparison with other programs revealed the validity of a new procedure. Tests with other data sets showed that BUGS can be used to efficiently fit animal model for medium sized data sets. Using these results quantitative geneticists can now easily use Bayesian approach to fit animal model in BUGS.

2009-03-11

Growth performance of station tested rams in Slovenia

This is our short communication for ASD 2009.

Update 2009-03-13: We had to shorten the manuscript to three pages. This of course lead to the exclusion of some results that will be published elsewhere.
Growth performance of station tested rams in Slovenia

2009-01-16

Življenjska prireja ovc bovške in oplemenjene bovške pasme

Pri diplomski nalogi Kendi Perčič (PDF 575 kB) smo analizirali življenjsko prirejo ovc bovške in oplemenjene bovške pasme. Sedaj smo en del tega dela pripravili za objavo v reviji - tokrat hrvaški. Prispevek si lahko ogledate tukaj.

2008-08-11

Calculation of PrP genotype and NSP type probabilities in Slovenian sheep

Paper "Calculation of PrP genotype and NSP type probabilities in Slovenian sheep" is finnaly finnished. I will send it to the journal today. This is a sort of a companion paper to the previous one where the methods were derived and tested, while in this paper we used those methods on all sheep breed in the Slovenian breeding programme for sheep. Here is the abstract:


PrP genotype probabilities in ungenotyped Slovenian sheep were calculated. Altogether 36,083 ewes and rams of various breeds were included into the analysis. PrP genotype was known for 10,504 animals. Five different PrP alleles were present in the data. Pedigree and genotype data structure differed between breeds. Iterative allelic peeling with incomplete penetrance model was used for the calculation of genotype probabilities for each animal given the genotype data of relatives. Analyses were performed for each breed separately. Additionally, NSP (National Scrapie Plan) type probabilities and the average NSP type were calculated from the genotype probabilities. Results were presented for live animals only. There were no animals with additionally identified PrP genotype or NSP type with certainty. With 95 % probability PrP genotype was additionally identified for 0.0 to 5.7 % animals of different breeds. NSP type was additionally identified with the same probability for 0.0 to 34.9 % animals of different breeds. We maintain that the low number of additional identifications was due to: a large number of alleles, intermediate allele frequencies, data structure, a uniform prior, and the use of incomplete penetrance model. Additional identifications provided some cost savings, but did not prove useful in the selection for scrapie resistance of the entire populations. The average NSP type should be used instead, since it can be calculated for all animals and encompasses all information from genotype probabilities.

2008-04-16

Use of genotype probabilities in selection on PrP genotype in sheep

I have finished paper on use of genotype probabilities in selection on PrP genotype in sheep. PrP genotype is one of the main factors for susceptibility to scrapie in sheep. Here is the link to the last version of the paper. And the abstract:

PrP genotype probabilities for ungenotyped animals of Jezersko-Solcava sheep breed were calculated. Data consisted of 10,429 animals among which 3,669 had PrP genotype data. There were 2,673 live ungenotyped animals. Five PrP haplotypes were present with the following frequencies: ARR 0.174, AHQ 0.074, ARH 0.083, ARQ 0.632, and VRQ 0.037. All 15 PrP genotypes were found. Iterative allelic peeling with incomplete penetrance model as implemented in GenoProb program was used for calculation of PrP genotype probabilities. There were only some additional PrP genotype and NSP type identifications with high probability. Main reasons for a low number of additional identifications can be attributed to large number of haplotypes with moderate frequencies, incomplete penetrance model, uniform prior, and inherently pedigree and genotype data structure. Incomplete penetrance model caused inflation of small probabilities but has proved to be very useful for field data where conflicting data rise due to pedigree of genotype errors. Novel parameters (maximal NSP type, average NSP type and its variance and accuracy) are proposed that make use of PrP genotype probabilities and can facilitate selection for scrapie resistance. Parameters were derived with emphasis on practical implementation of selection schemes based on NSP types. Maximal NSP type can be used to infer maximal potential scrapie susceptibility of individual ungenotyped animals as well as for whole flocks. Average NSP type makes full use of all PrP genotype probabilities and is the most useful and practical parameter for selection on NSP type and therefore PrP genotype. In addition, accuracy of average NSP type should also be used as selection criteria in order to assess the variability of average NSP type.

2007-12-28

Using Sweave with LyX

It has been a while since I managed to configure LyX (a great "word processor") to work with Sweave (literate programming using LaTeX and R - wikipedia page). The paper describing the whole thing will probably be published in the next issue of Rnews. Since, I received some inquiries, I have created a temporary web page at Google with all the files needed for the configuration, ...

Happy LyX-ing with R!