Showing posts with label PGS. Show all posts
Showing posts with label PGS. Show all posts

Thursday, December 19, 2024

'Doppelgänger' - a short story by Adam Carlton


---

“I don't want to die, Father,” I say to the priest, squirming on the plastic chair next to my bed.

Squirming.

Strange that a chaplain would be so squeamish. About heaven and hell, the baggage of his trade.

“This is the time to take stock of your life, to prepare for judgement, André... ,” the priest says.

But I’m no longer listening. I’m thirty-four, a mathematician at the École Polytechnique. Some success under my belt and so much more to give. I’m not minded to check out so easily.

“Thank you, Father,” I say, waving him away. He looks at me oddly, muttering something which sounds a lot like ‘denial’ under his breath.

So they send in the psychologist.


Dr P_ looks smart and businesslike. It’s pleasing not to be patronised for a change.

How is André feeling this morning? You are going to eat your porridge aren't you, or shall I be cross?

“Leukaemia, right? Terminal? How long do they give you?“

I decide to test him.

“It's a Poisson process. Mean survival time - as of today - is five days.”

He doesn’t blink an eye.

“Like radioactive decay, huh? We'd better get a move on, then!”

“If I were seventy,” I explain wearily, “with all my creative years behind me, well it would be different. I'd be happy to mulch back into the biosphere, my work done. But, ...”

I wave weakly and helplessly,

“... I'm just not ready to go.“

Dr P_ gives me a sympathetic look.

“Not much we can do about your personal Poisson process, I accept that. But reflect on this. Personal extinction is really something quite different. Something considerably more tractable.”

I may look confused; I certainly feel it.

The psychologist pulls up his tablet.

“I've got your details here. All those psychometric tests you did?”

He looks at me in mock admiration.

“IQ of 145, it says here. Not quite genius level but you must be one in a million.”

“1,300 in a million,” I correct him.

I'm beginning to warm to this guy.

“I'm looking at your Five-Factor and Myers-Briggs stats now,” he continues. “It says here you’re: flippant and facetious; careless about things which don’t interest you; tunnel-visioned and obsessive about the things which do.”

“That’s not so exceptional,” I say, “For a mathematician.”

He leans towards me, suddenly more serious.

“Personal identity is a strange old thing,” he says. “Every night you turn yourself off. Every morning you reboot yourself. A new you for a new day.”

He holds up his hand at my skeptical frown: where's he going with this?

“Listen. Hear me out. If you woke up tomorrow with amnesia, remembering nothing of your previous life, but still feeling some ineffable sense of your you-ness, is that still you?

I nod, humouring him. I know such things have happened. It wouldn't be ideal, but .. .

"Sure. It’s not like you’re dead."

Dr P_ stands up.

“Approximately a billion people in the world have been sequenced by now and their results put online for research purposes. In my professional capacity I have ..  André, are you listening?“

I suppress a yawn. Always interesting to hear about a colleague's work.

In any case, he hadn't stopped.

“I want you to think about the test-retest box around your psychometric score. Sure, it's tight but there are a lot of people out there. We'll run the PGS algorithms but I can tell you right now with greater than 99% certainty.  

"You have a doppelgänger out there somewhere."


Two days later a mathematics professor died at the tender age of thirty-four. It was remarked that his face had a peaceful expression, maybe even the hint of a smile.


It was another beautiful morning on the Ukrainian steppe as Katya left her parents’ cottage to cycle the ten kilometres to Institute 14, a school for the precociously-gifted.

On arrival, Katya logged-in and found an unexpected email. It was from Paris, France via anonymized routing. The author, a clinical psychologist, apologised for not being able to greet her properly - confidentiality required that he should know nothing of her identity or location in the world.

The message was simply for her information. A very promising mathematician, Professor André Z_ had just passed away, and the computers had identified her as being an uncanny match in personality and intellect.

There was a link to her mental twin's Wikipedia page. If she wished she could check out her doppelgänger’s life history and accomplishments. No specific action was required on her part and there would be no further communication.

Katya, mildly curious, scanned the article and then deleted the message. It was quickly forgotten as she hurried off to her first class of the day, in advanced statistics.


Tuesday, December 11, 2018

My mother in DNA.LAND

My mother, Beryl Seel, died on the 3rd of December 2015 aged 92. In August the previous year I had persuaded her to spit into the 23andMe sample collection tube and get sequenced.

"Now we can bring you back," I murmured, but she didn't seem keen.

I understand 23andMe pussyfooting around, not wishing to fall foul of the US regulators yet again. This means DNA.LAND if you want to play with polygenic scores. I uploaded her genome subset (the 23andMe downloaded text file) and here are some of the results.

As usual, click on an image to make it large enough to read.

---

This is the top-level screen showing some of the traits they can begin to predict

---

Educational attainment


My mother was predicted to have 14.7 years of schooling (two years at university). In fact she had 9, leaving school in 1937 as a 14 year old.

My prediction from DNA.LAND is also 14.7 years but in fact I had 19 years equivalent full-time education. Clare was predicted to have 14.4 years yet she is a graduate with an additional teacher-training qualification.

Heritability of EA is low and the small numbers of relevant SNPs available to DNA.LAND makes even the genetic component estimate a very noisy measure. See the discussion below.

---

Intelligence and IQ


This is marked as 'very preliminary' which - as they are quoting only 16 SNPs - seems about right.

Heritability of IQ is of the order of 80% but given this few genetic markers the signal will be swamped by the effects of all the other relevant SNPs which are unknown and whose contributions are therefore unmeasured.

About all you can say in favour of the exercise is that correlations between IQ-enhancing alleles based on an underlying soft sweep probably gives this slightly more predictive power than one might naively expect.

Note also we don't get a central IQ estimate with error bars. The normal distribution curve above is referenced to the DNA.LAND sign-up population, with unknown IQ parameters.

Anyway, FWIW, my mother seems to be one IQ point above the DNA.LAND average based on a super-noisy regression line.

---

Neuroticism



Neuroticism (emotionalism, anxiety) is one of the 'five factor' personality traits and is the opposite to emotional stability, calmness on this dimension. There are strong gender differences in this factor, women scoring as more emotional.

I have memories of a fair degree of maternal emotionalism, but perhaps that was par for the course.

---

Height



My DNA.LAND polygenic score was 182 cm, which is exactly right. My mother's (above) equates to 5' 7". That is taller than I remember, perhaps by an inch. I recall that growing up in the 1930s was not a picnic. My father was just under six foot.

My parents on their engagement: is she taller than average?

---

Discussion

How seriously should we take all this? Not very, at this point (except for height, where PGS scores are accurate predictors of genomic potential - although DNA.LAND aren't using the latest estimators).

Example regression line for scatter plot with correlation 0.38

The correlations between the small numbers of SNPs currently being used by DNA.LAND and the physical/cognitive variables they're trying to predict are very low. Educational attainment was discussed on their website as follows:
"Overall, educational attainment is estimated to have a heritability of around 20%. This means that in a population of individuals with varying years of education, 20% of that variation can be explained by variants in the individuals' genomes."
So the polygenic score will give the phenotype value on the regression line, but the true value could be a long way up or down depending on life events. My mother, a working class girl, left school at 14 for example. Nobody does that today.

Thousands of SNPs are implicated in complex polygenic traits like those listed above, yet relatively few are currently known (except for height) .. and DNA.LAND uses only those publicly available (a tiny subset). So their regression lines will be pretty inaccurate.

Finally, note that my mother is positioned with respect to DNA.LAND's user population which is self-selected. We don't know how the norms for this group compare to the general population.

It will get better, even if we're a way from the option of bringing her back.

Friday, November 23, 2018

Understanding Polygenic Scores (PGS)

In the years to come it will be very important to understand the concept of your polygenic score for traits such as height, weight, intelligence, personality and many others. In chapter 12 of his book, "Blueprint: How DNA Makes Us Who We Are" Robert Plomin gives a gentle introduction to the PGS concept which I excerpt here.

---

"Because polygenic scores are the basis for the DNA revolution in psychology, it is essential to understand what they are. A polygenic score is like any composite score that psychologists routinely use to create scales from items, such as those on a personality questionnaire. The goal of a polygenic score is to provide a single genetic index to predict a trait, whether schizophrenia, well-being or intelligence.

To get a concrete understanding of a polygenic score, consider a personality trait like shyness. A questionnaire to assess shyness includes multiple items in order to tap into different facets of shyness. For example, a typical shyness questionnaire will have items about how anxious you are in social situations and how much you avoid these situations for example, going to a party, meeting strangers and speaking up at a meeting. You might be asked to respond using a three-point scale (0 = not at all, 1 = sometimes, 2 = a lot).

A shyness score is created by adding these items, taking care to ‘reverse’ items as needed so that a high score means a high degree of shyness. If our shyness measure had ten items scored 0, 1 and 2, total scores could vary from 0 to 20. Simply adding the items like this treats each item as if it is equally useful, but all items are not equally useful. For this reason, items are often added after they are weighted by some criterion of their usefulness at capturing the construct of shyness.

This is exactly how polygenic scores are created, except that, instead of items on a questionnaire, we add up SNP genotypes. Like the three-point rating scale for shyness, SNP genotypes are scored as 0, 1 or 2, indicating the number of ‘increasing’ alleles, as in the example of the FTO SNP [a polymorphism implicated in weight gain].

In the same way that we can add up alleles for one SNP to create a genotypic score, we can also add up alleles for many SNPs to create a polygenic score, just as we add questionnaire items to create a shyness score.

The results from genome-wide association studies are used to select SNPs and to assign weights to each SNP. For example, in the GWA analysis of weight, the FTO SNP accounts for much more variance than other SNPs, so it should count for much more in a polygenic score for weight.

The following table shows how one individual’s polygenic score is created from ten SNPs. For the first SNP, this individual’s genotype is AT. For this SNP, the T allele happens to be the increasing allele that is positively associated with the trait. So, the individual’s genotypic score for this SNP is 1 because the genotype has only one increasing T allele.

Across the ten SNPs, the individual has a total of nine increasing alleles for the trait out of a possible score of 20. So, this individual would have a polygenic score just below the population average score of 10 for this trait.

This score merely adds the number of increasing alleles, which works reasonably well as a polygenic score.



However, we can increase its precision by weighting the genotypic score for each SNP by how much the SNP correlates with the trait. The correlation between each SNP and the trait is taken from the GWA analysis. If one SNP correlates five times more with the trait than another SNP such as SNP 1 versus SNP 10 it should count for five times as much in the polygenic score.

The weighted genotypic scores in the last column of the table are the product of the genotypic score for each SNP and the correlation with the trait. The sum of these weighted genotypic scores for the ten SNPs is 0.023.

This number isn’t as interpretable as the unweighted genotypic score of 9, which is just the sum of the ‘increasing’ alleles. However, both the unweighted polygenic score of 9 and the weighted score of 0.023 can be expressed simply as a percentile in the population. For this individual, both types would indicate a polygenic score just below average.

How many SNPs should go into a polygenic score? Initially, polygenic scores were created using only the genome-wide significant ‘hits’ from a GWA study. For weight, ninety-seven independent SNPs reached genome-wide significance. Creating a polygenic score from these top ninety-seven SNPs explains 1.2. percent of the variance in weight in independent samples. This is only slightly better than the prediction from the FTO SNP by itself, which explains 0.7 per cent of the variance.

Using only genome-wide significant hits is like demanding that each item in our shyness scale predicts significantly on its own. We don’t do this for other psychological scores because it is unrealistic to expect each item to stand on its own. The goal is to have a composite scale that is as useful as possible.

A better idea is to do what we do when we create other psychological scores: keep adding items as long as they add to the reliability and validity of the composite in independent samples. For polygenic scores, the key criterion is prediction. The new approach to polygenic scores is to keep adding SNPs as long as they add to the predictive power of the polygenic score in independent samples.

This is the strategy that has paid off in the last two years in producing powerful polygenic scores for psychological traits. Some false positives will be included in the polygenic score but that is acceptable as long as the signal increases relative to the noise, in the sense that the polygenic score predicts more variance.  ...

To interpret polygenic scores, it is important to keep in mind that they are always distributed like a bell-shaped curve, that is, a normal distribution. This bell-shaped curve is dictated by the fundamental law of probability, the central limit theorem, which is the basis for all statistics.




The normal distribution is found when many random events contribute to a phenomenon, like flipping a coin and counting the number of times the coin comes up heads. If you flip a coin ten times, you could get no heads or ten heads in a row, but most of the time the total number of heads will be between four and seven. If you do this many times, you will get a perfectly normal bell-shaped distribution, peaking at five, which will be the average number of heads. Flipping coins and counting heads is exactly analogous to counting the numbers of ‘increasing’ alleles from SNPs to construct polygenic scores for many individuals.

I will describe all my polygenic scores in terms of percentiles in the normal distribution. That is, to what extent is my polygenic score above or below the average polygenic score in the comparison sample, the 50th percentile?

It turns out that my polygenic score for height is at the 90th percentile. So, based on my DNA alone, knowing nothing else about me, you could predict that I am tall. And, in fact, I am 6 feet 5 inches. Of course, you can easily see that I am tall if you saw me, but with DNA you could tell that I am tall without even looking at me.

Most importantly, you could have predicted when I was born that I would be tall. Unlike any other predictors, polygenic scores are just as predictive from birth as from any other age because inherited DNA sequence does not change during life. In contrast, height at birth scarcely predicts adult height.

The predictive power of polygenic scores is greater than any other predictors, even the height of the individuals’ parents. Another advantage of polygenic scores over family resemblance is that parental height provides only a family-wide prediction that is the same for any child born to those parents.

In contrast, polygenic scores provide a prediction specific to each individual. In other words, my polygenic scores at birth would have predicted that I would be taller than expected on the basis of the average height of my parents.

Before looking at my other polygenic scores, one other general point needs to be highlighted about predicting individuals. My actual height is at the 99th percentile but my polygenic score is at the 90th percentile. Are polygenic scores sufficiently accurate for prediction?

For example, in TEDS [Twins Early Development Study - 1994 onwards], the polygenic score for height predicts 15 percent of the variance in actual height in these young adults. But 15 per cent is a long way from 100 per cent.

In fact, polygenic scores can never predict 100 per cent of the variance of any trait, because the ceiling for prediction is heritability. For height, heritability is 80 percent, but for psychological traits heritability is 50 percent, which means that polygenic score prediction is always going to be way south of perfect.




The big question is the extent to which polygenic scores will be able to predict all the heritable variance of traits. This gap is called missing heritability, and is described in the Notes section at the end of this book. "

--- [end of text extract] ---

Plomin's scatter plot looks rather messy but it hides the extent of clustering when you consider each decile separately.



As mentioned in the legend above, the vertical lines indicate the 95% confidence intervals. They clearly illustrate the linear trend line. Also, I suspect the limitations on sample size (20,000 pairs of twins in TEDS) for genetic studies.

For the top and bottom deciles, the following chart shows the extent of overlap.


Yet at the extremes the differences are very large. This is a general truth as regards traits whose values are are normally distributed.

Plomin finishes this section by emphasising yet again that due to the current lack of power (too few SNPs identified) and the ceiling of heritability, the PGS prediction is just that - a prediction with an error distribution around it. It is not deterministic.

I would comment that the error bars may be long right now, but as sample sizes get larger and non-additive effects are factored in, they can be made considerably smaller.

This is a small excerpt from an excellent book, by the way, which I reviewed here.