A.Turchin
http://www.scribd.com/doc/8553049/-
This is a list of articles about chaos theory, complexity theory, synergetics. If you want to see my real blogs please go to: http://www.0nothing1.blogspot.com/ it's in Russian, and: http://www.0dirtypurple1.blogspot.com/ it's in English -- some of my posts on Facebook. Это список статей о теории хаоса, теории сложности, синергетике. Если вы хотите увидеть мои настоящие блоги, перейдите к ссылкам выше.
четверг, 27 января 2011 г.
суббота, 22 января 2011 г.
The Truth Wears Off
Is there something wrong with the scientific method?by Jonah Lehrer
December 13, 2010
Read more http://www.newyorker.com/reporting/2010/12/13/101213fa_fact_lehrer#ixzz1Bl8DSHE5
The Truth Wears OffIs there something wrong with the scientific method?by Jonah Lehrer
December 13, 2010 Many results that are rigorously proved and accepted start shrinking in later studies.
Share Print E-Mail Single Page Keywords
Scientific Experiments; Decline Effect; Replicability; Scientists; Statistics; Jonathan Schooler; Scientific Theories On September 18, 2007, a few dozen neuroscientists, psychiatrists, and drug-company executives gathered in a hotel conference room in Brussels to hear some startling news. It had to do with a class of drugs known as atypical or second-generation antipsychotics, which came on the market in the early nineties. The drugs, sold under brand names such as Abilify, Seroquel, and Zyprexa, had been tested on schizophrenics in several large clinical trials, all of which had demonstrated a dramatic decrease in the subjects’ psychiatric symptoms. As a result, second-generation antipsychotics had become one of the fastest-growing and most profitable pharmaceutical classes. By 2001, Eli Lilly’s Zyprexa was generating more revenue than Prozac. It remains the company’s top-selling drug.
But the data presented at the Brussels meeting made it clear that something strange was happening: the therapeutic power of the drugs appeared to be steadily waning. A recent study showed an effect that was less than half of that documented in the first trials, in the early nineteen-nineties. Many researchers began to argue that the expensive pharmaceuticals weren’t any better than first-generation antipsychotics, which have been in use since the fifties. “In fact, sometimes they now look even worse,” John Davis, a professor of psychiatry at the University of Illinois at Chicago, told me.
Before the effectiveness of a drug can be confirmed, it must be tested and tested again. Different scientists in different labs need to repeat the protocols and publish their results. The test of replicability, as it’s known, is the foundation of modern research. Replicability is how the community enforces itself. It’s a safeguard for the creep of subjectivity. Most of the time, scientists know what results they want, and that can influence the results they get. The premise of replicability is that the scientific community can correct for these flaws.
But now all sorts of well-established, multiply confirmed findings have started to look increasingly uncertain. It’s as if our facts were losing their truth: claims that have been enshrined in textbooks are suddenly unprovable. This phenomenon doesn’t yet have an official name, but it’s occurring across a wide range of fields, from psychology to ecology. In the field of medicine, the phenomenon seems extremely widespread, affecting not only antipsychotics but also therapies ranging from cardiac stents to Vitamin E and antidepressants: Davis has a forthcoming analysis demonstrating that the efficacy of antidepressants has gone down as much as threefold in recent decades.
For many scientists, the effect is especially troubling because of what it exposes about the scientific process. If replication is what separates the rigor of science from the squishiness of pseudoscience, where do we put all these rigorously validated findings that can no longer be proved? Which results should we believe? Francis Bacon, the early-modern philosopher and pioneer of the scientific method, once declared that experiments were essential, because they allowed us to “put nature to the question.” But it appears that nature often gives us different answers.
from the issuecartoon banke-mail thisJonathan Schooler was a young graduate student at the University of Washington in the nineteen-eighties when he discovered a surprising new fact about language and memory. At the time, it was widely believed that the act of describing our memories improved them. But, in a series of clever experiments, Schooler demonstrated that subjects shown a face and asked to describe it were much less likely to recognize the face when shown it later than those who had simply looked at it. Schooler called the phenomenon “verbal overshadowing.”
The study turned him into an academic star. Since its initial publication, in 1990, it has been cited more than four hundred times. Before long, Schooler had extended the model to a variety of other tasks, such as remembering the taste of a wine, identifying the best strawberry jam, and solving difficult creative puzzles. In each instance, asking people to put their perceptions into words led to dramatic decreases in performance.
But while Schooler was publishing these results in highly reputable journals, a secret worry gnawed at him: it was proving difficult to replicate his earlier findings. “I’d often still see an effect, but the effect just wouldn’t be as strong,” he told me. “It was as if verbal overshadowing, my big new idea, was getting weaker.” At first, he assumed that he’d made an error in experimental design or a statistical miscalculation. But he couldn’t find anything wrong with his research. He then concluded that his initial batch of research subjects must have been unusually susceptible to verbal overshadowing. (John Davis, similarly, has speculated that part of the drop-off in the effectiveness of antipsychotics can be attributed to using subjects who suffer from milder forms of psychosis which are less likely to show dramatic improvement.) “It wasn’t a very satisfying explanation,” Schooler says. “One of my mentors told me that my real mistake was trying to replicate my work. He told me doing that was just setting myself up for disappointment.”
Schooler tried to put the problem out of his mind; his colleagues assured him that such things happened all the time. Over the next few years, he found new research questions, got married and had kids. But his replication problem kept on getting worse. His first attempt at replicating the 1990 study, in 1995, resulted in an effect that was thirty per cent smaller. The next year, the size of the effect shrank another thirty per cent. When other labs repeated Schooler’s experiments, they got a similar spread of data, with a distinct downward trend. “This was profoundly frustrating,” he says. “It was as if nature gave me this great result and then tried to take it back.” In private, Schooler began referring to the problem as “cosmic habituation,” by analogy to the decrease in response that occurs when individuals habituate to particular stimuli. “Habituation is why you don’t notice the stuff that’s always there,” Schooler says. “It’s an inevitable process of adjustment, a ratcheting down of excitement. I started joking that it was like the cosmos was habituating to my ideas. I took it very personally.”
Schooler is now a tenured professor at the University of California at Santa Barbara. He has curly black hair, pale-green eyes, and the relaxed demeanor of someone who lives five minutes away from his favorite beach. When he speaks, he tends to get distracted by his own digressions. He might begin with a point about memory, which reminds him of a favorite William James quote, which inspires a long soliloquy on the importance of introspection. Before long, we’re looking at pictures from Burning Man on his iPhone, which leads us back to the fragile nature of memory.
Although verbal overshadowing remains a widely accepted theory—it’s often invoked in the context of eyewitness testimony, for instance—Schooler is still a little peeved at the cosmos. “I know I should just move on already,” he says. “I really should stop talking about this. But I can’t.” That’s because he is convinced that he has stumbled on a serious problem, one that afflicts many of the most exciting new ideas in psychology.
One of the first demonstrations of this mysterious phenomenon came in the early nineteen-thirties. Joseph Banks Rhine, a psychologist at Duke, had developed an interest in the possibility of extrasensory perception, or E.S.P. Rhine devised an experiment featuring Zener cards, a special deck of twenty-five cards printed with one of five different symbols: a card was drawn from the deck and the subject was asked to guess the symbol. Most of Rhine’s subjects guessed about twenty per cent of the cards correctly, as you’d expect, but an undergraduate named Adam Linzmayer averaged nearly fifty per cent during his initial sessions, and pulled off several uncanny streaks, such as guessing nine cards in a row. The odds of this happening by chance are about one in two million. Linzmayer did it three times.
Rhine documented these stunning results in his notebook and prepared several papers for publication. But then, just as he began to believe in the possibility of extrasensory perception, the student lost his spooky talent. Between 1931 and 1933, Linzmayer guessed at the identity of another several thousand cards, but his success rate was now barely above chance. Rhine was forced to conclude that the student’s “extra-sensory perception ability has gone through a marked decline.” And Linzmayer wasn’t the only subject to experience such a drop-off: in nearly every case in which Rhine and others documented E.S.P. the effect dramatically diminished over time. Rhine called this trend the “decline effect.”
Schooler was fascinated by Rhine’s experimental struggles. Here was a scientist who had repeatedly documented the decline of his data; he seemed to have a talent for finding results that fell apart. In 2004, Schooler embarked on an ironic imitation of Rhine’s research: he tried to replicate this failure to replicate. In homage to Rhine’s interests, he decided to test for a parapsychological phenomenon known as precognition. The experiment itself was straightforward: he flashed a set of images to a subject and asked him or her to identify each one. Most of the time, the response was negative—the images were displayed too quickly to register. Then Schooler randomly selected half of the images to be shown again. What he wanted to know was whether the images that got a second showing were more likely to have been identified the first time around. Could subsequent exposure have somehow influenced the initial results? Could the effect become the cause?
from the issue
cartoon bank
e-mail this
The craziness of the hypothesis was the point: Schooler knows that precognition lacks a scientific explanation. But he wasn’t testing extrasensory powers; he was testing the decline effect. “At first, the data looked amazing, just as we’d expected,” Schooler says. “I couldn’t believe the amount of precognition we were finding. But then, as we kept on running subjects, the effect size”—a standard statistical measure—“kept on getting smaller and smaller.” The scientists eventually tested more than two thousand undergraduates. “In the end, our results looked just like Rhine’s,” Schooler said. “We found this strong paranormal effect, but it disappeared on us.”
The most likely explanation for the decline is an obvious one: regression to the mean. As the experiment is repeated, that is, an early statistical fluke gets cancelled out. The extrasensory powers of Schooler’s subjects didn’t decline—they were simply an illusion that vanished over time. And yet Schooler has noticed that many of the data sets that end up declining seem statistically solid—that is, they contain enough data that any regression to the mean shouldn’t be dramatic. “These are the results that pass all the tests,” he says. “The odds of them being random are typically quite remote, like one in a million. This means that the decline effect should almost never happen. But it happens all the time! Hell, it’s happened to me multiple times.” And this is why Schooler believes that the decline effect deserves more attention: its ubiquity seems to violate the laws of statistics. “Whenever I start talking about this, scientists get very nervous,” he says. “But I still want to know what happened to my results. Like most scientists, I assumed that it would get easier to document my effect over time. I’d get better at doing the experiments, at zeroing in on the conditions that produce verbal overshadowing. So why did the opposite happen? I’m convinced that we can use the tools of science to figure this out. First, though, we have to admit that we’ve got a problem.”
In 1991, the Danish zoologist Anders Møller, at Uppsala University, in Sweden, made a remarkable discovery about sex, barn swallows, and symmetry. It had long been known that the asymmetrical appearance of a creature was directly linked to the amount of mutation in its genome, so that more mutations led to more “fluctuating asymmetry.” (An easy way to measure asymmetry in humans is to compare the length of the fingers on each hand.) What Møller discovered is that female barn swallows were far more likely to mate with male birds that had long, symmetrical feathers. This suggested that the picky females were using symmetry as a proxy for the quality of male genes. Møller’s paper, which was published in Nature, set off a frenzy of research. Here was an easily measured, widely applicable indicator of genetic quality, and females could be shown to gravitate toward it. Aesthetics was really about genetics.
In the three years following, there were ten independent tests of the role of fluctuating asymmetry in sexual selection, and nine of them found a relationship between symmetry and male reproductive success. It didn’t matter if scientists were looking at the hairs on fruit flies or replicating the swallow studies—females seemed to prefer males with mirrored halves. Before long, the theory was applied to humans. Researchers found, for instance, that women preferred the smell of symmetrical men, but only during the fertile phase of the menstrual cycle. Other studies claimed that females had more orgasms when their partners were symmetrical, while a paper by anthropologists at Rutgers analyzed forty Jamaican dance routines and discovered that symmetrical men were consistently rated as better dancers.
Then the theory started to fall apart. In 1994, there were fourteen published tests of symmetry and sexual selection, and only eight found a correlation. In 1995, there were eight papers on the subject, and only four got a positive result. By 1998, when there were twelve additional investigations of fluctuating asymmetry, only a third of them confirmed the theory. Worse still, even the studies that yielded some positive result showed a steadily declining effect size. Between 1992 and 1997, the average effect size shrank by eighty per cent.
And it’s not just fluctuating asymmetry. In 2001, Michael Jennions, a biologist at the Australian National University, set out to analyze “temporal trends” across a wide range of subjects in ecology and evolutionary biology. He looked at hundreds of papers and forty-four meta-analyses (that is, statistical syntheses of related studies), and discovered a consistent decline effect over time, as many of the theories seemed to fade into irrelevance. In fact, even when numerous variables were controlled for—Jennions knew, for instance, that the same author might publish several critical papers, which could distort his analysis—there was still a significant decrease in the validity of the hypothesis, often within a year of publication. Jennions admits that his findings are troubling, but expresses a reluctance to talk about them publicly. “This is a very sensitive issue for scientists,” he says. “You know, we’re supposed to be dealing with hard facts, the stuff that’s supposed to stand the test of time. But when you see these trends you become a little more skeptical of things.”
What happened? Leigh Simmons, a biologist at the University of Western Australia, suggested one explanation when he told me about his initial enthusiasm for the theory: “I was really excited by fluctuating asymmetry. The early studies made the effect look very robust.” He decided to conduct a few experiments of his own, investigating symmetry in male horned beetles. “Unfortunately, I couldn’t find the effect,” he said. “But the worst part was that when I submitted these null results I had difficulty getting them published. The journals only wanted confirming data. It was too exciting an idea to disprove, at least back then.” For Simmons, the steep rise and slow fall of fluctuating asymmetry is a clear example of a scientific paradigm, one of those intellectual fads that both guide and constrain research: after a new paradigm is proposed, the peer-review process is tilted toward positive results. But then, after a few years, the academic incentives shift—the paradigm has become entrenched—so that the most notable results are now those that disprove the theory.
from the issue
cartoon bank
e-mail this
Jennions, similarly, argues that the decline effect is largely a product of publication bias, or the tendency of scientists and scientific journals to prefer positive data over null results, which is what happens when no effect is found. The bias was first identified by the statistician Theodore Sterling, in 1959, after he noticed that ninety-seven per cent of all published psychological studies with statistically significant data found the effect they were looking for. A “significant” result is defined as any data point that would be produced by chance less than five per cent of the time. This ubiquitous test was invented in 1922 by the English mathematician Ronald Fisher, who picked five per cent as the boundary line, somewhat arbitrarily, because it made pencil and slide-rule calculations easier. Sterling saw that if ninety-seven per cent of psychology studies were proving their hypotheses, either psychologists were extraordinarily lucky or they published only the outcomes of successful experiments. In recent years, publication bias has mostly been seen as a problem for clinical trials, since pharmaceutical companies are less interested in publishing results that aren’t favorable. But it’s becoming increasingly clear that publication bias also produces major distortions in fields without large corporate incentives, such as psychology and ecology.
While publication bias almost certainly plays a role in the decline effect, it remains an incomplete explanation. For one thing, it fails to account for the initial prevalence of positive results among studies that never even get submitted to journals. It also fails to explain the experience of people like Schooler, who have been unable to replicate their initial data despite their best efforts. Richard Palmer, a biologist at the University of Alberta, who has studied the problems surrounding fluctuating asymmetry, suspects that an equally significant issue is the selective reporting of results—the data that scientists choose to document in the first place. Palmer’s most convincing evidence relies on a statistical tool known as a funnel graph. When a large number of studies have been done on a single subject, the data should follow a pattern: studies with a large sample size should all cluster around a common value—the true result—whereas those with a smaller sample size should exhibit a random scattering, since they’re subject to greater sampling error. This pattern gives the graph its name, since the distribution resembles a funnel.
The funnel graph visually captures the distortions of selective reporting. For instance, after Palmer plotted every study of fluctuating asymmetry, he noticed that the distribution of results with smaller sample sizes wasn’t random at all but instead skewed heavily toward positive results. Palmer has since documented a similar problem in several other contested subject areas. “Once I realized that selective reporting is everywhere in science, I got quite depressed,” Palmer told me. “As a researcher, you’re always aware that there might be some nonrandom patterns, but I had no idea how widespread it is.” In a recent review article, Palmer summarized the impact of selective reporting on his field: “We cannot escape the troubling conclusion that some—perhaps many—cherished generalities are at best exaggerated in their biological significance and at worst a collective illusion nurtured by strong a-priori beliefs often repeated.”
Palmer emphasizes that selective reporting is not the same as scientific fraud. Rather, the problem seems to be one of subtle omissions and unconscious misperceptions, as researchers struggle to make sense of their results. Stephen Jay Gould referred to this as the “shoehorning” process. “A lot of scientific measurement is really hard,” Simmons told me. “If you’re talking about fluctuating asymmetry, then it’s a matter of minuscule differences between the right and left sides of an animal. It’s millimetres of a tail feather. And so maybe a researcher knows that he’s measuring a good male”—an animal that has successfully mated—“and he knows that it’s supposed to be symmetrical. Well, that act of measurement is going to be vulnerable to all sorts of perception biases. That’s not a cynical statement. That’s just the way human beings work.”
One of the classic examples of selective reporting concerns the testing of acupuncture in different countries. While acupuncture is widely accepted as a medical treatment in various Asian countries, its use is much more contested in the West. These cultural differences have profoundly influenced the results of clinical trials. Between 1966 and 1995, there were forty-seven studies of acupuncture in China, Taiwan, and Japan, and every single trial concluded that acupuncture was an effective treatment. During the same period, there were ninety-four clinical trials of acupuncture in the United States, Sweden, and the U.K., and only fifty-six per cent of these studies found any therapeutic benefits. As Palmer notes, this wide discrepancy suggests that scientists find ways to confirm their preferred hypothesis, disregarding what they don’t want to see. Our beliefs are a form of blindness.
John Ioannidis, an epidemiologist at Stanford University, argues that such distortions are a serious issue in biomedical research. “These exaggerations are why the decline has become so common,” he says. “It’d be really great if the initial studies gave us an accurate summary of things. But they don’t. And so what happens is we waste a lot of money treating millions of patients and doing lots of follow-up studies on other themes based on results that are misleading.” In 2005, Ioannidis published an article in the Journal of the American Medical Association that looked at the forty-nine most cited clinical-research studies in three major medical journals. Forty-five of these studies reported positive results, suggesting that the intervention being tested was effective. Because most of these studies were randomized controlled trials—the “gold standard” of medical evidence—they tended to have a significant impact on clinical practice, and led to the spread of treatments such as hormone replacement therapy for menopausal women and daily low-dose aspirin to prevent heart attacks and strokes. Nevertheless, the data Ioannidis found were disturbing: of the thirty-four claims that had been subject to replication, forty-one per cent had either been directly contradicted or had their effect sizes significantly downgraded.
The situation is even worse when a subject is fashionable. In recent years, for instance, there have been hundreds of studies on the various genes that control the differences in disease risk between men and women. These findings have included everything from the mutations responsible for the increased risk of schizophrenia to the genes underlying hypertension. Ioannidis and his colleagues looked at four hundred and thirty-two of these claims. They quickly discovered that the vast majority had serious flaws. But the most troubling fact emerged when he looked at the test of replication: out of four hundred and thirty-two claims, only a single one was consistently replicable. “This doesn’t mean that none of these claims will turn out to be true,” he says. “But, given that most of them were done badly, I wouldn’t hold my breath.”
from the issue
cartoon bank
e-mail this
According to Ioannidis, the main problem is that too many researchers engage in what he calls “significance chasing,” or finding ways to interpret the data so that it passes the statistical test of significance—the ninety-five-per-cent boundary invented by Ronald Fisher. “The scientists are so eager to pass this magical test that they start playing around with the numbers, trying to find anything that seems worthy,” Ioannidis says. In recent years, Ioannidis has become increasingly blunt about the pervasiveness of the problem. One of his most cited papers has a deliberately provocative title: “Why Most Published Research Findings Are False.”
The problem of selective reporting is rooted in a fundamental cognitive flaw, which is that we like proving ourselves right and hate being wrong. “It feels good to validate a hypothesis,” Ioannidis said. “It feels even better when you’ve got a financial interest in the idea or your career depends upon it. And that’s why, even after a claim has been systematically disproven”—he cites, for instance, the early work on hormone replacement therapy, or claims involving various vitamins—“you still see some stubborn researchers citing the first few studies that show a strong effect. They really want to believe that it’s true.”
That’s why Schooler argues that scientists need to become more rigorous about data collection before they publish. “We’re wasting too much time chasing after bad studies and underpowered experiments,” he says. The current “obsession” with replicability distracts from the real problem, which is faulty design. He notes that nobody even tries to replicate most science papers—there are simply too many. (According to Nature, a third of all studies never even get cited, let alone repeated.) “I’ve learned the hard way to be exceedingly careful,” Schooler says. “Every researcher should have to spell out, in advance, how many subjects they’re going to use, and what exactly they’re testing, and what constitutes a sufficient level of proof. We have the tools to be much more transparent about our experiments.”
In a forthcoming paper, Schooler recommends the establishment of an open-source database, in which researchers are required to outline their planned investigations and document all their results. “I think this would provide a huge increase in access to scientific work and give us a much better way to judge the quality of an experiment,” Schooler says. “It would help us finally deal with all these issues that the decline effect is exposing.”
Although such reforms would mitigate the dangers of publication bias and selective reporting, they still wouldn’t erase the decline effect. This is largely because scientific research will always be shadowed by a force that can’t be curbed, only contained: sheer randomness. Although little research has been done on the experimental dangers of chance and happenstance, the research that exists isn’t encouraging.
In the late nineteen-nineties, John Crabbe, a neuroscientist at the Oregon Health and Science University, conducted an experiment that showed how unknowable chance events can skew tests of replicability. He performed a series of experiments on mouse behavior in three different science labs: in Albany, New York; Edmonton, Alberta; and Portland, Oregon. Before he conducted the experiments, he tried to standardize every variable he could think of. The same strains of mice were used in each lab, shipped on the same day from the same supplier. The animals were raised in the same kind of enclosure, with the same brand of sawdust bedding. They had been exposed to the same amount of incandescent light, were living with the same number of littermates, and were fed the exact same type of chow pellets. When the mice were handled, it was with the same kind of surgical glove, and when they were tested it was on the same equipment, at the same time in the morning.
The premise of this test of replicability, of course, is that each of the labs should have generated the same pattern of results. “If any set of experiments should have passed the test, it should have been ours,” Crabbe says. “But that’s not the way it turned out.” In one experiment, Crabbe injected a particular strain of mouse with cocaine. In Portland the mice given the drug moved, on average, six hundred centimetres more than they normally did; in Albany they moved seven hundred and one additional centimetres. But in the Edmonton lab they moved more than five thousand additional centimetres. Similar deviations were observed in a test of anxiety. Furthermore, these inconsistencies didn’t follow any detectable pattern. In Portland one strain of mouse proved most anxious, while in Albany another strain won that distinction.
The disturbing implication of the Crabbe study is that a lot of extraordinary scientific data are nothing but noise. The hyperactivity of those coked-up Edmonton mice wasn’t an interesting new fact—it was a meaningless outlier, a by-product of invisible variables we don’t understand. The problem, of course, is that such dramatic findings are also the most likely to get published in prestigious journals, since the data are both statistically significant and entirely unexpected. Grants get written, follow-up studies are conducted. The end result is a scientific accident that can take years to unravel.
from the issue
cartoon bank
e-mail this
This suggests that the decline effect is actually a decline of illusion. While Karl Popper imagined falsification occurring with a single, definitive experiment—Galileo refuted Aristotelian mechanics in an afternoon—the process turns out to be much messier than that. Many scientific theories continue to be considered true even after failing numerous experimental tests. Verbal overshadowing might exhibit the decline effect, but it remains extensively relied upon within the field. The same holds for any number of phenomena, from the disappearing benefits of second-generation antipsychotics to the weak coupling ratio exhibited by decaying neutrons, which appears to have fallen by more than ten standard deviations between 1969 and 2001. Even the law of gravity hasn’t always been perfect at predicting real-world phenomena. (In one test, physicists measuring gravity by means of deep boreholes in the Nevada desert found a two-and-a-half-per-cent discrepancy between the theoretical predictions and the actual data.) Despite these findings, second-generation antipsychotics are still widely prescribed, and our model of the neutron hasn’t changed. The law of gravity remains the same.
Such anomalies demonstrate the slipperiness of empiricism. Although many scientific ideas generate conflicting results and suffer from falling effect sizes, they continue to get cited in the textbooks and drive standard medical practice. Why? Because these ideas seem true. Because they make sense. Because we can’t bear to let them go. And this is why the decline effect is so troubling. Not because it reveals the human fallibility of science, in which data are tweaked and beliefs shape perceptions. (Such shortcomings aren’t surprising, at least for scientists.) And not because it reveals that many of our most exciting theories are fleeting fads and will soon be rejected. (That idea has been around since Thomas Kuhn.) The decline effect is troubling because it reminds us how difficult it is to prove anything. We like to pretend that our experiments define the truth for us. But that’s often not the case. Just because an idea is true doesn’t mean it can be proved. And just because an idea can be proved doesn’t mean it’s true. When the experiments are done, we still have to choose what to believe. ♦
Read more http://www.newyorker.com/reporting/2010/12/13/101213fa_fact_lehrer#ixzz1BlB1sWp0
Is there something wrong with the scientific method?by Jonah Lehrer
December 13, 2010
Read more http://www.newyorker.com/reporting/2010/12/13/101213fa_fact_lehrer#ixzz1Bl8DSHE5
The Truth Wears OffIs there something wrong with the scientific method?by Jonah Lehrer
December 13, 2010 Many results that are rigorously proved and accepted start shrinking in later studies.
Share Print E-Mail Single Page Keywords
Scientific Experiments; Decline Effect; Replicability; Scientists; Statistics; Jonathan Schooler; Scientific Theories On September 18, 2007, a few dozen neuroscientists, psychiatrists, and drug-company executives gathered in a hotel conference room in Brussels to hear some startling news. It had to do with a class of drugs known as atypical or second-generation antipsychotics, which came on the market in the early nineties. The drugs, sold under brand names such as Abilify, Seroquel, and Zyprexa, had been tested on schizophrenics in several large clinical trials, all of which had demonstrated a dramatic decrease in the subjects’ psychiatric symptoms. As a result, second-generation antipsychotics had become one of the fastest-growing and most profitable pharmaceutical classes. By 2001, Eli Lilly’s Zyprexa was generating more revenue than Prozac. It remains the company’s top-selling drug.
But the data presented at the Brussels meeting made it clear that something strange was happening: the therapeutic power of the drugs appeared to be steadily waning. A recent study showed an effect that was less than half of that documented in the first trials, in the early nineteen-nineties. Many researchers began to argue that the expensive pharmaceuticals weren’t any better than first-generation antipsychotics, which have been in use since the fifties. “In fact, sometimes they now look even worse,” John Davis, a professor of psychiatry at the University of Illinois at Chicago, told me.
Before the effectiveness of a drug can be confirmed, it must be tested and tested again. Different scientists in different labs need to repeat the protocols and publish their results. The test of replicability, as it’s known, is the foundation of modern research. Replicability is how the community enforces itself. It’s a safeguard for the creep of subjectivity. Most of the time, scientists know what results they want, and that can influence the results they get. The premise of replicability is that the scientific community can correct for these flaws.
But now all sorts of well-established, multiply confirmed findings have started to look increasingly uncertain. It’s as if our facts were losing their truth: claims that have been enshrined in textbooks are suddenly unprovable. This phenomenon doesn’t yet have an official name, but it’s occurring across a wide range of fields, from psychology to ecology. In the field of medicine, the phenomenon seems extremely widespread, affecting not only antipsychotics but also therapies ranging from cardiac stents to Vitamin E and antidepressants: Davis has a forthcoming analysis demonstrating that the efficacy of antidepressants has gone down as much as threefold in recent decades.
For many scientists, the effect is especially troubling because of what it exposes about the scientific process. If replication is what separates the rigor of science from the squishiness of pseudoscience, where do we put all these rigorously validated findings that can no longer be proved? Which results should we believe? Francis Bacon, the early-modern philosopher and pioneer of the scientific method, once declared that experiments were essential, because they allowed us to “put nature to the question.” But it appears that nature often gives us different answers.
from the issuecartoon banke-mail thisJonathan Schooler was a young graduate student at the University of Washington in the nineteen-eighties when he discovered a surprising new fact about language and memory. At the time, it was widely believed that the act of describing our memories improved them. But, in a series of clever experiments, Schooler demonstrated that subjects shown a face and asked to describe it were much less likely to recognize the face when shown it later than those who had simply looked at it. Schooler called the phenomenon “verbal overshadowing.”
The study turned him into an academic star. Since its initial publication, in 1990, it has been cited more than four hundred times. Before long, Schooler had extended the model to a variety of other tasks, such as remembering the taste of a wine, identifying the best strawberry jam, and solving difficult creative puzzles. In each instance, asking people to put their perceptions into words led to dramatic decreases in performance.
But while Schooler was publishing these results in highly reputable journals, a secret worry gnawed at him: it was proving difficult to replicate his earlier findings. “I’d often still see an effect, but the effect just wouldn’t be as strong,” he told me. “It was as if verbal overshadowing, my big new idea, was getting weaker.” At first, he assumed that he’d made an error in experimental design or a statistical miscalculation. But he couldn’t find anything wrong with his research. He then concluded that his initial batch of research subjects must have been unusually susceptible to verbal overshadowing. (John Davis, similarly, has speculated that part of the drop-off in the effectiveness of antipsychotics can be attributed to using subjects who suffer from milder forms of psychosis which are less likely to show dramatic improvement.) “It wasn’t a very satisfying explanation,” Schooler says. “One of my mentors told me that my real mistake was trying to replicate my work. He told me doing that was just setting myself up for disappointment.”
Schooler tried to put the problem out of his mind; his colleagues assured him that such things happened all the time. Over the next few years, he found new research questions, got married and had kids. But his replication problem kept on getting worse. His first attempt at replicating the 1990 study, in 1995, resulted in an effect that was thirty per cent smaller. The next year, the size of the effect shrank another thirty per cent. When other labs repeated Schooler’s experiments, they got a similar spread of data, with a distinct downward trend. “This was profoundly frustrating,” he says. “It was as if nature gave me this great result and then tried to take it back.” In private, Schooler began referring to the problem as “cosmic habituation,” by analogy to the decrease in response that occurs when individuals habituate to particular stimuli. “Habituation is why you don’t notice the stuff that’s always there,” Schooler says. “It’s an inevitable process of adjustment, a ratcheting down of excitement. I started joking that it was like the cosmos was habituating to my ideas. I took it very personally.”
Schooler is now a tenured professor at the University of California at Santa Barbara. He has curly black hair, pale-green eyes, and the relaxed demeanor of someone who lives five minutes away from his favorite beach. When he speaks, he tends to get distracted by his own digressions. He might begin with a point about memory, which reminds him of a favorite William James quote, which inspires a long soliloquy on the importance of introspection. Before long, we’re looking at pictures from Burning Man on his iPhone, which leads us back to the fragile nature of memory.
Although verbal overshadowing remains a widely accepted theory—it’s often invoked in the context of eyewitness testimony, for instance—Schooler is still a little peeved at the cosmos. “I know I should just move on already,” he says. “I really should stop talking about this. But I can’t.” That’s because he is convinced that he has stumbled on a serious problem, one that afflicts many of the most exciting new ideas in psychology.
One of the first demonstrations of this mysterious phenomenon came in the early nineteen-thirties. Joseph Banks Rhine, a psychologist at Duke, had developed an interest in the possibility of extrasensory perception, or E.S.P. Rhine devised an experiment featuring Zener cards, a special deck of twenty-five cards printed with one of five different symbols: a card was drawn from the deck and the subject was asked to guess the symbol. Most of Rhine’s subjects guessed about twenty per cent of the cards correctly, as you’d expect, but an undergraduate named Adam Linzmayer averaged nearly fifty per cent during his initial sessions, and pulled off several uncanny streaks, such as guessing nine cards in a row. The odds of this happening by chance are about one in two million. Linzmayer did it three times.
Rhine documented these stunning results in his notebook and prepared several papers for publication. But then, just as he began to believe in the possibility of extrasensory perception, the student lost his spooky talent. Between 1931 and 1933, Linzmayer guessed at the identity of another several thousand cards, but his success rate was now barely above chance. Rhine was forced to conclude that the student’s “extra-sensory perception ability has gone through a marked decline.” And Linzmayer wasn’t the only subject to experience such a drop-off: in nearly every case in which Rhine and others documented E.S.P. the effect dramatically diminished over time. Rhine called this trend the “decline effect.”
Schooler was fascinated by Rhine’s experimental struggles. Here was a scientist who had repeatedly documented the decline of his data; he seemed to have a talent for finding results that fell apart. In 2004, Schooler embarked on an ironic imitation of Rhine’s research: he tried to replicate this failure to replicate. In homage to Rhine’s interests, he decided to test for a parapsychological phenomenon known as precognition. The experiment itself was straightforward: he flashed a set of images to a subject and asked him or her to identify each one. Most of the time, the response was negative—the images were displayed too quickly to register. Then Schooler randomly selected half of the images to be shown again. What he wanted to know was whether the images that got a second showing were more likely to have been identified the first time around. Could subsequent exposure have somehow influenced the initial results? Could the effect become the cause?
from the issue
cartoon bank
e-mail this
The craziness of the hypothesis was the point: Schooler knows that precognition lacks a scientific explanation. But he wasn’t testing extrasensory powers; he was testing the decline effect. “At first, the data looked amazing, just as we’d expected,” Schooler says. “I couldn’t believe the amount of precognition we were finding. But then, as we kept on running subjects, the effect size”—a standard statistical measure—“kept on getting smaller and smaller.” The scientists eventually tested more than two thousand undergraduates. “In the end, our results looked just like Rhine’s,” Schooler said. “We found this strong paranormal effect, but it disappeared on us.”
The most likely explanation for the decline is an obvious one: regression to the mean. As the experiment is repeated, that is, an early statistical fluke gets cancelled out. The extrasensory powers of Schooler’s subjects didn’t decline—they were simply an illusion that vanished over time. And yet Schooler has noticed that many of the data sets that end up declining seem statistically solid—that is, they contain enough data that any regression to the mean shouldn’t be dramatic. “These are the results that pass all the tests,” he says. “The odds of them being random are typically quite remote, like one in a million. This means that the decline effect should almost never happen. But it happens all the time! Hell, it’s happened to me multiple times.” And this is why Schooler believes that the decline effect deserves more attention: its ubiquity seems to violate the laws of statistics. “Whenever I start talking about this, scientists get very nervous,” he says. “But I still want to know what happened to my results. Like most scientists, I assumed that it would get easier to document my effect over time. I’d get better at doing the experiments, at zeroing in on the conditions that produce verbal overshadowing. So why did the opposite happen? I’m convinced that we can use the tools of science to figure this out. First, though, we have to admit that we’ve got a problem.”
In 1991, the Danish zoologist Anders Møller, at Uppsala University, in Sweden, made a remarkable discovery about sex, barn swallows, and symmetry. It had long been known that the asymmetrical appearance of a creature was directly linked to the amount of mutation in its genome, so that more mutations led to more “fluctuating asymmetry.” (An easy way to measure asymmetry in humans is to compare the length of the fingers on each hand.) What Møller discovered is that female barn swallows were far more likely to mate with male birds that had long, symmetrical feathers. This suggested that the picky females were using symmetry as a proxy for the quality of male genes. Møller’s paper, which was published in Nature, set off a frenzy of research. Here was an easily measured, widely applicable indicator of genetic quality, and females could be shown to gravitate toward it. Aesthetics was really about genetics.
In the three years following, there were ten independent tests of the role of fluctuating asymmetry in sexual selection, and nine of them found a relationship between symmetry and male reproductive success. It didn’t matter if scientists were looking at the hairs on fruit flies or replicating the swallow studies—females seemed to prefer males with mirrored halves. Before long, the theory was applied to humans. Researchers found, for instance, that women preferred the smell of symmetrical men, but only during the fertile phase of the menstrual cycle. Other studies claimed that females had more orgasms when their partners were symmetrical, while a paper by anthropologists at Rutgers analyzed forty Jamaican dance routines and discovered that symmetrical men were consistently rated as better dancers.
Then the theory started to fall apart. In 1994, there were fourteen published tests of symmetry and sexual selection, and only eight found a correlation. In 1995, there were eight papers on the subject, and only four got a positive result. By 1998, when there were twelve additional investigations of fluctuating asymmetry, only a third of them confirmed the theory. Worse still, even the studies that yielded some positive result showed a steadily declining effect size. Between 1992 and 1997, the average effect size shrank by eighty per cent.
And it’s not just fluctuating asymmetry. In 2001, Michael Jennions, a biologist at the Australian National University, set out to analyze “temporal trends” across a wide range of subjects in ecology and evolutionary biology. He looked at hundreds of papers and forty-four meta-analyses (that is, statistical syntheses of related studies), and discovered a consistent decline effect over time, as many of the theories seemed to fade into irrelevance. In fact, even when numerous variables were controlled for—Jennions knew, for instance, that the same author might publish several critical papers, which could distort his analysis—there was still a significant decrease in the validity of the hypothesis, often within a year of publication. Jennions admits that his findings are troubling, but expresses a reluctance to talk about them publicly. “This is a very sensitive issue for scientists,” he says. “You know, we’re supposed to be dealing with hard facts, the stuff that’s supposed to stand the test of time. But when you see these trends you become a little more skeptical of things.”
What happened? Leigh Simmons, a biologist at the University of Western Australia, suggested one explanation when he told me about his initial enthusiasm for the theory: “I was really excited by fluctuating asymmetry. The early studies made the effect look very robust.” He decided to conduct a few experiments of his own, investigating symmetry in male horned beetles. “Unfortunately, I couldn’t find the effect,” he said. “But the worst part was that when I submitted these null results I had difficulty getting them published. The journals only wanted confirming data. It was too exciting an idea to disprove, at least back then.” For Simmons, the steep rise and slow fall of fluctuating asymmetry is a clear example of a scientific paradigm, one of those intellectual fads that both guide and constrain research: after a new paradigm is proposed, the peer-review process is tilted toward positive results. But then, after a few years, the academic incentives shift—the paradigm has become entrenched—so that the most notable results are now those that disprove the theory.
from the issue
cartoon bank
e-mail this
Jennions, similarly, argues that the decline effect is largely a product of publication bias, or the tendency of scientists and scientific journals to prefer positive data over null results, which is what happens when no effect is found. The bias was first identified by the statistician Theodore Sterling, in 1959, after he noticed that ninety-seven per cent of all published psychological studies with statistically significant data found the effect they were looking for. A “significant” result is defined as any data point that would be produced by chance less than five per cent of the time. This ubiquitous test was invented in 1922 by the English mathematician Ronald Fisher, who picked five per cent as the boundary line, somewhat arbitrarily, because it made pencil and slide-rule calculations easier. Sterling saw that if ninety-seven per cent of psychology studies were proving their hypotheses, either psychologists were extraordinarily lucky or they published only the outcomes of successful experiments. In recent years, publication bias has mostly been seen as a problem for clinical trials, since pharmaceutical companies are less interested in publishing results that aren’t favorable. But it’s becoming increasingly clear that publication bias also produces major distortions in fields without large corporate incentives, such as psychology and ecology.
While publication bias almost certainly plays a role in the decline effect, it remains an incomplete explanation. For one thing, it fails to account for the initial prevalence of positive results among studies that never even get submitted to journals. It also fails to explain the experience of people like Schooler, who have been unable to replicate their initial data despite their best efforts. Richard Palmer, a biologist at the University of Alberta, who has studied the problems surrounding fluctuating asymmetry, suspects that an equally significant issue is the selective reporting of results—the data that scientists choose to document in the first place. Palmer’s most convincing evidence relies on a statistical tool known as a funnel graph. When a large number of studies have been done on a single subject, the data should follow a pattern: studies with a large sample size should all cluster around a common value—the true result—whereas those with a smaller sample size should exhibit a random scattering, since they’re subject to greater sampling error. This pattern gives the graph its name, since the distribution resembles a funnel.
The funnel graph visually captures the distortions of selective reporting. For instance, after Palmer plotted every study of fluctuating asymmetry, he noticed that the distribution of results with smaller sample sizes wasn’t random at all but instead skewed heavily toward positive results. Palmer has since documented a similar problem in several other contested subject areas. “Once I realized that selective reporting is everywhere in science, I got quite depressed,” Palmer told me. “As a researcher, you’re always aware that there might be some nonrandom patterns, but I had no idea how widespread it is.” In a recent review article, Palmer summarized the impact of selective reporting on his field: “We cannot escape the troubling conclusion that some—perhaps many—cherished generalities are at best exaggerated in their biological significance and at worst a collective illusion nurtured by strong a-priori beliefs often repeated.”
Palmer emphasizes that selective reporting is not the same as scientific fraud. Rather, the problem seems to be one of subtle omissions and unconscious misperceptions, as researchers struggle to make sense of their results. Stephen Jay Gould referred to this as the “shoehorning” process. “A lot of scientific measurement is really hard,” Simmons told me. “If you’re talking about fluctuating asymmetry, then it’s a matter of minuscule differences between the right and left sides of an animal. It’s millimetres of a tail feather. And so maybe a researcher knows that he’s measuring a good male”—an animal that has successfully mated—“and he knows that it’s supposed to be symmetrical. Well, that act of measurement is going to be vulnerable to all sorts of perception biases. That’s not a cynical statement. That’s just the way human beings work.”
One of the classic examples of selective reporting concerns the testing of acupuncture in different countries. While acupuncture is widely accepted as a medical treatment in various Asian countries, its use is much more contested in the West. These cultural differences have profoundly influenced the results of clinical trials. Between 1966 and 1995, there were forty-seven studies of acupuncture in China, Taiwan, and Japan, and every single trial concluded that acupuncture was an effective treatment. During the same period, there were ninety-four clinical trials of acupuncture in the United States, Sweden, and the U.K., and only fifty-six per cent of these studies found any therapeutic benefits. As Palmer notes, this wide discrepancy suggests that scientists find ways to confirm their preferred hypothesis, disregarding what they don’t want to see. Our beliefs are a form of blindness.
John Ioannidis, an epidemiologist at Stanford University, argues that such distortions are a serious issue in biomedical research. “These exaggerations are why the decline has become so common,” he says. “It’d be really great if the initial studies gave us an accurate summary of things. But they don’t. And so what happens is we waste a lot of money treating millions of patients and doing lots of follow-up studies on other themes based on results that are misleading.” In 2005, Ioannidis published an article in the Journal of the American Medical Association that looked at the forty-nine most cited clinical-research studies in three major medical journals. Forty-five of these studies reported positive results, suggesting that the intervention being tested was effective. Because most of these studies were randomized controlled trials—the “gold standard” of medical evidence—they tended to have a significant impact on clinical practice, and led to the spread of treatments such as hormone replacement therapy for menopausal women and daily low-dose aspirin to prevent heart attacks and strokes. Nevertheless, the data Ioannidis found were disturbing: of the thirty-four claims that had been subject to replication, forty-one per cent had either been directly contradicted or had their effect sizes significantly downgraded.
The situation is even worse when a subject is fashionable. In recent years, for instance, there have been hundreds of studies on the various genes that control the differences in disease risk between men and women. These findings have included everything from the mutations responsible for the increased risk of schizophrenia to the genes underlying hypertension. Ioannidis and his colleagues looked at four hundred and thirty-two of these claims. They quickly discovered that the vast majority had serious flaws. But the most troubling fact emerged when he looked at the test of replication: out of four hundred and thirty-two claims, only a single one was consistently replicable. “This doesn’t mean that none of these claims will turn out to be true,” he says. “But, given that most of them were done badly, I wouldn’t hold my breath.”
from the issue
cartoon bank
e-mail this
According to Ioannidis, the main problem is that too many researchers engage in what he calls “significance chasing,” or finding ways to interpret the data so that it passes the statistical test of significance—the ninety-five-per-cent boundary invented by Ronald Fisher. “The scientists are so eager to pass this magical test that they start playing around with the numbers, trying to find anything that seems worthy,” Ioannidis says. In recent years, Ioannidis has become increasingly blunt about the pervasiveness of the problem. One of his most cited papers has a deliberately provocative title: “Why Most Published Research Findings Are False.”
The problem of selective reporting is rooted in a fundamental cognitive flaw, which is that we like proving ourselves right and hate being wrong. “It feels good to validate a hypothesis,” Ioannidis said. “It feels even better when you’ve got a financial interest in the idea or your career depends upon it. And that’s why, even after a claim has been systematically disproven”—he cites, for instance, the early work on hormone replacement therapy, or claims involving various vitamins—“you still see some stubborn researchers citing the first few studies that show a strong effect. They really want to believe that it’s true.”
That’s why Schooler argues that scientists need to become more rigorous about data collection before they publish. “We’re wasting too much time chasing after bad studies and underpowered experiments,” he says. The current “obsession” with replicability distracts from the real problem, which is faulty design. He notes that nobody even tries to replicate most science papers—there are simply too many. (According to Nature, a third of all studies never even get cited, let alone repeated.) “I’ve learned the hard way to be exceedingly careful,” Schooler says. “Every researcher should have to spell out, in advance, how many subjects they’re going to use, and what exactly they’re testing, and what constitutes a sufficient level of proof. We have the tools to be much more transparent about our experiments.”
In a forthcoming paper, Schooler recommends the establishment of an open-source database, in which researchers are required to outline their planned investigations and document all their results. “I think this would provide a huge increase in access to scientific work and give us a much better way to judge the quality of an experiment,” Schooler says. “It would help us finally deal with all these issues that the decline effect is exposing.”
Although such reforms would mitigate the dangers of publication bias and selective reporting, they still wouldn’t erase the decline effect. This is largely because scientific research will always be shadowed by a force that can’t be curbed, only contained: sheer randomness. Although little research has been done on the experimental dangers of chance and happenstance, the research that exists isn’t encouraging.
In the late nineteen-nineties, John Crabbe, a neuroscientist at the Oregon Health and Science University, conducted an experiment that showed how unknowable chance events can skew tests of replicability. He performed a series of experiments on mouse behavior in three different science labs: in Albany, New York; Edmonton, Alberta; and Portland, Oregon. Before he conducted the experiments, he tried to standardize every variable he could think of. The same strains of mice were used in each lab, shipped on the same day from the same supplier. The animals were raised in the same kind of enclosure, with the same brand of sawdust bedding. They had been exposed to the same amount of incandescent light, were living with the same number of littermates, and were fed the exact same type of chow pellets. When the mice were handled, it was with the same kind of surgical glove, and when they were tested it was on the same equipment, at the same time in the morning.
The premise of this test of replicability, of course, is that each of the labs should have generated the same pattern of results. “If any set of experiments should have passed the test, it should have been ours,” Crabbe says. “But that’s not the way it turned out.” In one experiment, Crabbe injected a particular strain of mouse with cocaine. In Portland the mice given the drug moved, on average, six hundred centimetres more than they normally did; in Albany they moved seven hundred and one additional centimetres. But in the Edmonton lab they moved more than five thousand additional centimetres. Similar deviations were observed in a test of anxiety. Furthermore, these inconsistencies didn’t follow any detectable pattern. In Portland one strain of mouse proved most anxious, while in Albany another strain won that distinction.
The disturbing implication of the Crabbe study is that a lot of extraordinary scientific data are nothing but noise. The hyperactivity of those coked-up Edmonton mice wasn’t an interesting new fact—it was a meaningless outlier, a by-product of invisible variables we don’t understand. The problem, of course, is that such dramatic findings are also the most likely to get published in prestigious journals, since the data are both statistically significant and entirely unexpected. Grants get written, follow-up studies are conducted. The end result is a scientific accident that can take years to unravel.
from the issue
cartoon bank
e-mail this
This suggests that the decline effect is actually a decline of illusion. While Karl Popper imagined falsification occurring with a single, definitive experiment—Galileo refuted Aristotelian mechanics in an afternoon—the process turns out to be much messier than that. Many scientific theories continue to be considered true even after failing numerous experimental tests. Verbal overshadowing might exhibit the decline effect, but it remains extensively relied upon within the field. The same holds for any number of phenomena, from the disappearing benefits of second-generation antipsychotics to the weak coupling ratio exhibited by decaying neutrons, which appears to have fallen by more than ten standard deviations between 1969 and 2001. Even the law of gravity hasn’t always been perfect at predicting real-world phenomena. (In one test, physicists measuring gravity by means of deep boreholes in the Nevada desert found a two-and-a-half-per-cent discrepancy between the theoretical predictions and the actual data.) Despite these findings, second-generation antipsychotics are still widely prescribed, and our model of the neutron hasn’t changed. The law of gravity remains the same.
Such anomalies demonstrate the slipperiness of empiricism. Although many scientific ideas generate conflicting results and suffer from falling effect sizes, they continue to get cited in the textbooks and drive standard medical practice. Why? Because these ideas seem true. Because they make sense. Because we can’t bear to let them go. And this is why the decline effect is so troubling. Not because it reveals the human fallibility of science, in which data are tweaked and beliefs shape perceptions. (Such shortcomings aren’t surprising, at least for scientists.) And not because it reveals that many of our most exciting theories are fleeting fads and will soon be rejected. (That idea has been around since Thomas Kuhn.) The decline effect is troubling because it reminds us how difficult it is to prove anything. We like to pretend that our experiments define the truth for us. But that’s often not the case. Just because an idea is true doesn’t mean it can be proved. And just because an idea can be proved doesn’t mean it’s true. When the experiments are done, we still have to choose what to believe. ♦
Read more http://www.newyorker.com/reporting/2010/12/13/101213fa_fact_lehrer#ixzz1BlB1sWp0
среда, 8 декабря 2010 г.
Как глобальные корпорации стимулируют рост национальных стартапов
00:03 РБК daily
Центральная тема в разговорах о модернизации российской экономики — спрос на инновации, предъявляемый крупным бизнесом. В частности, нередко говорится, что технологический спрос со стороны больших корпораций является причиной рождения множества малых стартапов. Это не совсем точно — малые инновационные компании рождаются по другим причинам. Однако «большие» и «маленькие» являются элементами одной системы, которая устанавливает между ними тесные взаимосвязи. Если эта система связей сформирована — экономика развивается по инновационному сценарию, если нет — число успешных стартапов остается минимальным.
Условие развития любой малой компании — неординарный подход. Или продукт необычен, или создана новаторская бизнес-модель, или система продаж — нетрадиционна. Иными словами, у малой компании должно быть новое конкурентное преимущество — в противном случае она не выйдет на рынок, плотно занятый другими. Допустим, я хочу выйти на рынок бензоколонок. Это категорически невозможно, если я не предложу инновационного решения — например, создам технологию, при которой человеку не нужно выходить из машины. Архаичный вариант — люди заливают вам топливо, инновационный вариант — машина подъезжает, пистолет соединяется с баком, а оплата осуществляется с мобильного телефона. Остается вопрос с ценой: сколько готов заплатить потребитель за дополнительное удобство? Если рынок принимает новый продукт, и я очень точно, определяю оптимальную премию за новизну — у меня появляется шанс выиграть. Таким образом, причина рождения стартапов — высокая конкуренция. Большие корпорации уже заняли рынок, а малым еще предстоит это сделать. Флеш-карты, Windows Media Player, CD-ROM-привод — тысячи продуктов родились в стартапах. Но мы как потребители познакомились с ними через большие корпорации. Почему же большой бизнес поставил на поток продукты, изобретенные в стартапах?
Существует всем известный (и желанный многими) путь развития стартапа. Стартап, как мы рассмотрели на примере бензоколонок, разрабатывает востребованную рынком технологию и предлагает потребителю продукт этой технологии по «правильной» цене — но это лишь предпосылки для успеха. Дальше вам надо как можно быстрее выходить к инвестору, получать от него деньги, чтобы опять же максимально быстро выскочить на рынок и занять на нем максимально широкую нишу по объему продаж. Все это делается, в общем, с одной единственной целью — верхняя строка отчета о прибылях и убытках. Инвестиционное сообщество на основании этой строки, собственно, и определяет будущую стоимость акций (я немного утрирую, но в целом это так). Таким образом, добившись высокого объема продаж, стартап готов для проведения IPO по максимально высокой цене — отличный exit (выход из инвестиций) для любого инвестора. Классическая модель, только в ней очень много «если». Если венчурный инвестор вовремя не найдется, или не получится быстро достичь достаточного объема продаж, то пока вы пробиваете железные двери, крупная компания может наладить серийное производство, оставив вас вне игры. Но давайте взглянем на ситуацию по-другому.
Необязательно создавать новый рынок или новую нишу на традиционном — можно встраиваться в ниши, созданные другими. Каждая крупная корпорация создает новые ниши десятками, если не сотнями. Преимущество большой компании не в том, что она инновационна, а в том, что она обеспечена каналами продаж. Я хорошо знаком с бизнесом одной большой компании — знаю, в частности, что она аккумулировала cash в размере 16 млрд долл. Компания занимается высокотехнологичным производством, развивает НИОКР. Так почему бы, казалось бы, не вложить эти деньги в исследования и разработки? Проблема в том, что как только компания заявит о таком вложении, цена на ее акции существенно упадет, ведь НИОКР — вещь рискованная: то ли получится, то ли нет. Таким образом, все крупные корпорации (и это не преувеличение) «сидят» на наличных располагая очень ограниченным набором решений, как этими деньгами распорядиться. Возвращаясь к компании, которую я упомянул в связи с 16 млрд, отмечу, что ее бюджет на собственный НИОКР составляет порядка 2 млрд долл. и примерно в пять раз больше вкладывается в покупку стартапов. И любой крупный высокотехнологичный бизнес тратит на поиск подходящих для покупки активов колоссальное время, силы и средства (Google, без сомнения, станет крупнейшим игроком M&A в 2010 году).
Я уверен, что сотрудничество с большими корпорациями — естественный путь развития российских стартапов. Задача бизнес-инкубаторов и технопарков, как я ее вижу, заключается в том, чтобы корпорации четко сформулировали, в чем заключаются их потребности в области инноваций, и донесли эту информацию до стартапов. А с другой стороны, нужно найти стартапы, способные удовлетворить спрос большого бизнеса, и вырастить их до того состояния, когда они смогут с этой задачей успешно справиться. На глобальном рынке все корпорации заявляют об интересе к инновационным стартапам — на рынке России им еще предстоит этот интерес продекларировать.
Илья Толстов, генеральный директор технопарка «Ингрия»
Читать полностью: http://www.rbcdaily.ru/2010/12/08/focus/562949979263403.shtml
http://www.rbcdaily.ru/2010/12/08/focus/562949979263403
Центральная тема в разговорах о модернизации российской экономики — спрос на инновации, предъявляемый крупным бизнесом. В частности, нередко говорится, что технологический спрос со стороны больших корпораций является причиной рождения множества малых стартапов. Это не совсем точно — малые инновационные компании рождаются по другим причинам. Однако «большие» и «маленькие» являются элементами одной системы, которая устанавливает между ними тесные взаимосвязи. Если эта система связей сформирована — экономика развивается по инновационному сценарию, если нет — число успешных стартапов остается минимальным.
Условие развития любой малой компании — неординарный подход. Или продукт необычен, или создана новаторская бизнес-модель, или система продаж — нетрадиционна. Иными словами, у малой компании должно быть новое конкурентное преимущество — в противном случае она не выйдет на рынок, плотно занятый другими. Допустим, я хочу выйти на рынок бензоколонок. Это категорически невозможно, если я не предложу инновационного решения — например, создам технологию, при которой человеку не нужно выходить из машины. Архаичный вариант — люди заливают вам топливо, инновационный вариант — машина подъезжает, пистолет соединяется с баком, а оплата осуществляется с мобильного телефона. Остается вопрос с ценой: сколько готов заплатить потребитель за дополнительное удобство? Если рынок принимает новый продукт, и я очень точно, определяю оптимальную премию за новизну — у меня появляется шанс выиграть. Таким образом, причина рождения стартапов — высокая конкуренция. Большие корпорации уже заняли рынок, а малым еще предстоит это сделать. Флеш-карты, Windows Media Player, CD-ROM-привод — тысячи продуктов родились в стартапах. Но мы как потребители познакомились с ними через большие корпорации. Почему же большой бизнес поставил на поток продукты, изобретенные в стартапах?
Существует всем известный (и желанный многими) путь развития стартапа. Стартап, как мы рассмотрели на примере бензоколонок, разрабатывает востребованную рынком технологию и предлагает потребителю продукт этой технологии по «правильной» цене — но это лишь предпосылки для успеха. Дальше вам надо как можно быстрее выходить к инвестору, получать от него деньги, чтобы опять же максимально быстро выскочить на рынок и занять на нем максимально широкую нишу по объему продаж. Все это делается, в общем, с одной единственной целью — верхняя строка отчета о прибылях и убытках. Инвестиционное сообщество на основании этой строки, собственно, и определяет будущую стоимость акций (я немного утрирую, но в целом это так). Таким образом, добившись высокого объема продаж, стартап готов для проведения IPO по максимально высокой цене — отличный exit (выход из инвестиций) для любого инвестора. Классическая модель, только в ней очень много «если». Если венчурный инвестор вовремя не найдется, или не получится быстро достичь достаточного объема продаж, то пока вы пробиваете железные двери, крупная компания может наладить серийное производство, оставив вас вне игры. Но давайте взглянем на ситуацию по-другому.
Необязательно создавать новый рынок или новую нишу на традиционном — можно встраиваться в ниши, созданные другими. Каждая крупная корпорация создает новые ниши десятками, если не сотнями. Преимущество большой компании не в том, что она инновационна, а в том, что она обеспечена каналами продаж. Я хорошо знаком с бизнесом одной большой компании — знаю, в частности, что она аккумулировала cash в размере 16 млрд долл. Компания занимается высокотехнологичным производством, развивает НИОКР. Так почему бы, казалось бы, не вложить эти деньги в исследования и разработки? Проблема в том, что как только компания заявит о таком вложении, цена на ее акции существенно упадет, ведь НИОКР — вещь рискованная: то ли получится, то ли нет. Таким образом, все крупные корпорации (и это не преувеличение) «сидят» на наличных располагая очень ограниченным набором решений, как этими деньгами распорядиться. Возвращаясь к компании, которую я упомянул в связи с 16 млрд, отмечу, что ее бюджет на собственный НИОКР составляет порядка 2 млрд долл. и примерно в пять раз больше вкладывается в покупку стартапов. И любой крупный высокотехнологичный бизнес тратит на поиск подходящих для покупки активов колоссальное время, силы и средства (Google, без сомнения, станет крупнейшим игроком M&A в 2010 году).
Я уверен, что сотрудничество с большими корпорациями — естественный путь развития российских стартапов. Задача бизнес-инкубаторов и технопарков, как я ее вижу, заключается в том, чтобы корпорации четко сформулировали, в чем заключаются их потребности в области инноваций, и донесли эту информацию до стартапов. А с другой стороны, нужно найти стартапы, способные удовлетворить спрос большого бизнеса, и вырастить их до того состояния, когда они смогут с этой задачей успешно справиться. На глобальном рынке все корпорации заявляют об интересе к инновационным стартапам — на рынке России им еще предстоит этот интерес продекларировать.
Илья Толстов, генеральный директор технопарка «Ингрия»
Читать полностью: http://www.rbcdaily.ru/2010/12/08/focus/562949979263403.shtml
http://www.rbcdaily.ru/2010/12/08/focus/562949979263403
пятница, 1 октября 2010 г.
Can Economic Risk Be Tamed And Markets 'Know'? Increasingly: No.
by Stuart Kauffman
In this post, I want to show that we often do not know what can happen in the technological evolution of the economy, hence in the most fundamental sense, in general, risk can neither be tamed, nor can markets accurately factor that risk in. As technological innovation accelerates, this problem is likely to become ever worse. This challenges foundational ideas in economics, and gives the lie to the noisome cant of some professional Conservative think tanks and noisy talking heads that we must leave everything to unregulated free markets that somehow always "know".
In my previous blog, Breaking the Galilean Spell, I described Darwinian "exaptations" in biological evolution. As I describe below, the same ideas apply in technological evolution. Exaptations are Darwin's idea that a causal property of an organism, like heart sounds, that was not the selected function of the heart, (to pump blood), and of no selective use in the current environment, could become of selective significance in a different environment, so be selected. He, and other biologists since him, including myself, note that typically a new function arises in the biosphere.
I told of the evolution of the swim bladder, that adjusts neutral buoyancy in the water column by the ratio of air to water in the bladder, by exaptation from the lungs of lung fish. Water got into the lungs of some fish, and this organ was poised to evolve into a swim bladder. I asked: 1) Did a new function come to exist in the biosphere? Yes, neutral buoyancy. 2) Did this new function alter the future course of evolution? Yes, new species, proteins, and niches. So the becoming of the world was changed. 3). Can we prestate all possible exaptations, just for humans? We all seem to agree the answer is "NO". I've asked thousands of people. And I noted that parts of why the answer is no, were: How would we prestate all possible selective conditions? How would we know we had listed them all? How would we prestate the features of one or many organisms that might become preadaptations? There seems no way to do these things.
Then I defined the"Actual" and the "Adjacent Possible" of a litre of 1000 initial kinds of small molecules. Call these 1000 the "Actual". The Adjacent Possible are those new kinds of molecules that could form in single reaction steps from the Actual. Then, by the paragraph above, we cannot prestate all the possibilities of the Adjacent Possible evolution of the biosphere. It follows that we do not know all of what can happen in evolution.
Further we cannot even make probability statements about the evolution of the biosphere by such exaptations, for we do not know the set of all the possibilities, called the "sample space", of the evolutionary process. Not knowing the sample space, we cannot construct a probability "measure".
Most startlingly, if a natural law is a compact description of the regularities of a process, we cannot have a "sufficient" natural law for the emergence of swim bladders. The becoming of the universe is not fully describable by natural law - thus "breaking the Galilean Spell" since Galileo and Newton that all that unfolds in the universe is describable by natural law.
Then what about the technological evolution of the economy? The same ideas apply. Once again, we often do not know what can happen, so, in general, risk cannot be tamed, and, in general, markets cannot always "know". More, the problem is getting worse.
I tell a story I am told is true, of exaptations in the econosphere. Some engineers were trying to invent the tractor. They knew they needed a very large engine, hence a very large engine block. They mounted the block on successively larger chassis. All broke in turn. Finally one engineer said, "You know, the engine block is so big and rigid, we can hang everything off the engine block and use it as the chassis. It worked! And that is how tractors are made. So too were formula racing cars for some time.
Now the use of the rigidity of the engine block as the chassis by the engineers is a Darwinian exaptation - it is the use of an unused causal feature of a system for a novel function. Did a new function arise? Yes, tractors. Could we prestate all uses of an engine block, or screw driver for that matter? No. Who knows what novel uses an engine block or screw driver might find? How would we list them all, know we had listed them all, or what causal features of the engine block or screw driver might be of use for some purpose? Again, there seems to be no way to do this.
Technological evolution is full of Darwinian exaptations. Most inventions are not used for their initial purpose. A blog ago,my colleagues and I discussed the evolution of the early computer, invented to calculate shell trajectories in WWII, that enabled the invention of the Apple personal computer, which in turn afforded the opportunities to successively invent: word processing, hence Microsoft, files, sharing files between buildings at CERN, the world wide web, eBay sales on the web, Google search engines on the web, and, at last, Facebook. Thomas Watson Sr. foresaw use for three computers in the 1940s at IBM. Well, no, Thomas Watson. Did any of us foresee Facebook or Google 25 years ago? No.
Once again, like the biosphere's evolution, we cannot prestate technological evolution. Once again, not only do we not know what will happen, often we really do not know what can happen.
I tell a funny and true personal story. A number of years ago, a baby Bell sought my advice about investing $2 billion for fiber optic cables. They would pay me $5000.00. I reasoned closely: "I don't know anything about fiber optic cables...But $5000.00 is $5000.00." I agreed and learned a lot in a day about fiber optic cables.
At the end of the day, from nowhere, an image came to me and I said, "How do you know some kid won't learn to keep empty tin cans in the air around the globe with beebee guns and bounce signals around the world? Your fiber optic cables would be worthless!" Seven pairs of eyes glimmered at my stupidity. The company invested the $2,000,000,000 in fiber optic cables. Six years or so later, satellites were launched that bounced telephone signals around the globe, rendering those cables useless for a long time.
I am, of course, unreasonably proud of my tin cans, but the question is: Was the baby Bell company stupid? NO. They could not know what would happen.
Now economists often think they can tame "risk" and that markets typically can and do correctly factor in risk. Suppose the baby Bell had issued bonds backed by expected revenues from the fiber optic cables, in order to buy the cables. Suppose Moody rated them, (at the same time Moody was payed by the baby Bell to do the rating, with the obvious conflict of interest), and rated the bonds AA. Maybe hundreds of thousands of the proverbial little old ladies invested in the bonds and they proved worthless.
Were the baby Bell and Moodys able to tame risk? No. Did the market "know"? No.
There are at least two different reasons risk might not be tamed. The first concerns what are called "power law distributions" with no means or variances. Power laws are distributions plotting the logarithm of, say the number of Nile floods of a certain size on the X axis, and the logarithm of the frequency of floods of each size on the Y axis. It is mathematically true that if the slope of this line is flatter than -1.0, the distribution has no average, or mean, nor higher moments like variance. So no amount of data can tell you what might happen. Others have made this point, for example Taleb in "The Black Swan".
What I am talking about is much more radical. In the case above, at least we knew what variable to measure: the size distribution of Nile floods. But in the case I am considering, the baby Bell, fiber optic cables and the unforeseen invention of satellites to bounce telephone signals around the globe, this lethal risk was not previsible, and we did not even know what variables to measure. This is the "unknown unknown".
So for Taleb's reason, and I think more deeply, for my reason, we cannot, in general, tame risk.
Economists want to believe that fancy trading algorithms, or the market itself, can price in risk accurately. In general, it cannot. We do not know before hand what variables to measure. Economists will treat such innovations as "exogenous shocks", but they are not exogenous. The invention of satellites to bounce signals around the globe grew organically out of technological evolution which, as Brian Arthur says in "The Nature of Technology" always grows out of existing technology. Very often this growth into the Adjacent Possible of the econosphere cannot be prestated.
Our pace of entry into the technological adjacent possible is accelerating, and with it the frequency of our encounters with the unknown unknown. With this accelerating pace, taming risk is ever more beyond reach. Free markets cannot always "know" and are likely to know ever less adequately as we explode into the Adjacent Possible technologically.
In this post, I want to show that we often do not know what can happen in the technological evolution of the economy, hence in the most fundamental sense, in general, risk can neither be tamed, nor can markets accurately factor that risk in. As technological innovation accelerates, this problem is likely to become ever worse. This challenges foundational ideas in economics, and gives the lie to the noisome cant of some professional Conservative think tanks and noisy talking heads that we must leave everything to unregulated free markets that somehow always "know".
In my previous blog, Breaking the Galilean Spell, I described Darwinian "exaptations" in biological evolution. As I describe below, the same ideas apply in technological evolution. Exaptations are Darwin's idea that a causal property of an organism, like heart sounds, that was not the selected function of the heart, (to pump blood), and of no selective use in the current environment, could become of selective significance in a different environment, so be selected. He, and other biologists since him, including myself, note that typically a new function arises in the biosphere.
I told of the evolution of the swim bladder, that adjusts neutral buoyancy in the water column by the ratio of air to water in the bladder, by exaptation from the lungs of lung fish. Water got into the lungs of some fish, and this organ was poised to evolve into a swim bladder. I asked: 1) Did a new function come to exist in the biosphere? Yes, neutral buoyancy. 2) Did this new function alter the future course of evolution? Yes, new species, proteins, and niches. So the becoming of the world was changed. 3). Can we prestate all possible exaptations, just for humans? We all seem to agree the answer is "NO". I've asked thousands of people. And I noted that parts of why the answer is no, were: How would we prestate all possible selective conditions? How would we know we had listed them all? How would we prestate the features of one or many organisms that might become preadaptations? There seems no way to do these things.
Then I defined the"Actual" and the "Adjacent Possible" of a litre of 1000 initial kinds of small molecules. Call these 1000 the "Actual". The Adjacent Possible are those new kinds of molecules that could form in single reaction steps from the Actual. Then, by the paragraph above, we cannot prestate all the possibilities of the Adjacent Possible evolution of the biosphere. It follows that we do not know all of what can happen in evolution.
Further we cannot even make probability statements about the evolution of the biosphere by such exaptations, for we do not know the set of all the possibilities, called the "sample space", of the evolutionary process. Not knowing the sample space, we cannot construct a probability "measure".
Most startlingly, if a natural law is a compact description of the regularities of a process, we cannot have a "sufficient" natural law for the emergence of swim bladders. The becoming of the universe is not fully describable by natural law - thus "breaking the Galilean Spell" since Galileo and Newton that all that unfolds in the universe is describable by natural law.
Then what about the technological evolution of the economy? The same ideas apply. Once again, we often do not know what can happen, so, in general, risk cannot be tamed, and, in general, markets cannot always "know". More, the problem is getting worse.
I tell a story I am told is true, of exaptations in the econosphere. Some engineers were trying to invent the tractor. They knew they needed a very large engine, hence a very large engine block. They mounted the block on successively larger chassis. All broke in turn. Finally one engineer said, "You know, the engine block is so big and rigid, we can hang everything off the engine block and use it as the chassis. It worked! And that is how tractors are made. So too were formula racing cars for some time.
Now the use of the rigidity of the engine block as the chassis by the engineers is a Darwinian exaptation - it is the use of an unused causal feature of a system for a novel function. Did a new function arise? Yes, tractors. Could we prestate all uses of an engine block, or screw driver for that matter? No. Who knows what novel uses an engine block or screw driver might find? How would we list them all, know we had listed them all, or what causal features of the engine block or screw driver might be of use for some purpose? Again, there seems to be no way to do this.
Technological evolution is full of Darwinian exaptations. Most inventions are not used for their initial purpose. A blog ago,my colleagues and I discussed the evolution of the early computer, invented to calculate shell trajectories in WWII, that enabled the invention of the Apple personal computer, which in turn afforded the opportunities to successively invent: word processing, hence Microsoft, files, sharing files between buildings at CERN, the world wide web, eBay sales on the web, Google search engines on the web, and, at last, Facebook. Thomas Watson Sr. foresaw use for three computers in the 1940s at IBM. Well, no, Thomas Watson. Did any of us foresee Facebook or Google 25 years ago? No.
Once again, like the biosphere's evolution, we cannot prestate technological evolution. Once again, not only do we not know what will happen, often we really do not know what can happen.
I tell a funny and true personal story. A number of years ago, a baby Bell sought my advice about investing $2 billion for fiber optic cables. They would pay me $5000.00. I reasoned closely: "I don't know anything about fiber optic cables...But $5000.00 is $5000.00." I agreed and learned a lot in a day about fiber optic cables.
At the end of the day, from nowhere, an image came to me and I said, "How do you know some kid won't learn to keep empty tin cans in the air around the globe with beebee guns and bounce signals around the world? Your fiber optic cables would be worthless!" Seven pairs of eyes glimmered at my stupidity. The company invested the $2,000,000,000 in fiber optic cables. Six years or so later, satellites were launched that bounced telephone signals around the globe, rendering those cables useless for a long time.
I am, of course, unreasonably proud of my tin cans, but the question is: Was the baby Bell company stupid? NO. They could not know what would happen.
Now economists often think they can tame "risk" and that markets typically can and do correctly factor in risk. Suppose the baby Bell had issued bonds backed by expected revenues from the fiber optic cables, in order to buy the cables. Suppose Moody rated them, (at the same time Moody was payed by the baby Bell to do the rating, with the obvious conflict of interest), and rated the bonds AA. Maybe hundreds of thousands of the proverbial little old ladies invested in the bonds and they proved worthless.
Were the baby Bell and Moodys able to tame risk? No. Did the market "know"? No.
There are at least two different reasons risk might not be tamed. The first concerns what are called "power law distributions" with no means or variances. Power laws are distributions plotting the logarithm of, say the number of Nile floods of a certain size on the X axis, and the logarithm of the frequency of floods of each size on the Y axis. It is mathematically true that if the slope of this line is flatter than -1.0, the distribution has no average, or mean, nor higher moments like variance. So no amount of data can tell you what might happen. Others have made this point, for example Taleb in "The Black Swan".
What I am talking about is much more radical. In the case above, at least we knew what variable to measure: the size distribution of Nile floods. But in the case I am considering, the baby Bell, fiber optic cables and the unforeseen invention of satellites to bounce telephone signals around the globe, this lethal risk was not previsible, and we did not even know what variables to measure. This is the "unknown unknown".
So for Taleb's reason, and I think more deeply, for my reason, we cannot, in general, tame risk.
Economists want to believe that fancy trading algorithms, or the market itself, can price in risk accurately. In general, it cannot. We do not know before hand what variables to measure. Economists will treat such innovations as "exogenous shocks", but they are not exogenous. The invention of satellites to bounce signals around the globe grew organically out of technological evolution which, as Brian Arthur says in "The Nature of Technology" always grows out of existing technology. Very often this growth into the Adjacent Possible of the econosphere cannot be prestated.
Our pace of entry into the technological adjacent possible is accelerating, and with it the frequency of our encounters with the unknown unknown. With this accelerating pace, taming risk is ever more beyond reach. Free markets cannot always "know" and are likely to know ever less adequately as we explode into the Adjacent Possible technologically.
Beyond The 'Washington Consensus:' Economic Webs And Growth
by Stuart Kauffman
Enlarge Romeo Gacad/AFP/Getty Images
Do we need to go beyond the "Washington Consensus" to spur global economic growth?
Romeo Gacad/AFP/Getty Images
Do we need to go beyond the "Washington Consensus" to spur global economic growth?
A body of economic theory known as the "Washington Consensus" has guided attempts to spur global economic growth for over two decades. This Consensus has largely failed, and contemporary economic growth theory seems mostly at a loss to understand the failure.
Meanwhile, the disparity of income between rich and poor countries has grown from 4-to-1 in Adam Smith's time to 70-to-1 now. No one knows why.
Part of the Washington Consensus is based on the work of economist David Ricardo and his theory of national competitive advantage. Ethiopia is good at producing coffee, Alberta is good at producing wheat. The two should trade freely to their joint advantage. Unfortunately, this leaves both as what we call "sub-critical" economies that cannot endogenously generate a growing diversity of goods and production capacities, hence wealth and growth.
We propose an extension of growth theory based on the structure of economic webs and their sub-critical versus supra-critical expansion into an un-prestatable "Adjacent Possible" of the "econosphere" that may be of significant help in guiding attempts to to drive growth.
Four of us, myself, Naresh Singh, acting vice president of the Canadian Partnership Branch of the Canadian International Development Agency, Rohinton Medhora, vice president of programs at the International Development Research Centre in Canada and Ricardo Hausmann, of the Harvard Kennedy School and a past Finance Minister for Venezuela, have co-authored this post and hope, with others, to develop new economic growth theory beyond the Washington Consensus to guide practical economic aid on the ground.
Growth and development strategies are recognized to be in disarray by an increasing number of institutions and stakeholders. It used to be believed that short to-do lists were enough to guide countries to prosperity, whether they be the peace, easy taxes and a tolerable administration of justice of Adam Smith or the openness, sound money and contract enforcement of Larry Summers.
In 1990, John Williamson christened the "Washington Consensus." It advocated fiscal and monetary soundness, openness to trade and investment, financial liberalization and regulation, privatization, deregulation and secure property rights. It was boiled down to a 10-item policy checklist for governments to follow.
After two decades of major reforms in many countries that followed these principles — with generous support from multilateral and bilateral agencies — the results are, in general, surprising and disappointing. The star reformers have not been the star performers and no evidence has been found that these policies promoted growth, once extremely bad policy outcomes are excluded from the sample.
Explaining the disparity in wealth across countries was the question that led Adam Smith to write The Wealth of Nations in 1776. But at that time the income differences to be explained were of the order of 4-to-1. In the meantime, they have grown to over 70-to-1, despite knowledge of the to-do lists and billions spent on aid.
Much of the conventional approach to development strategies is based on the idea that growth and development happen naturally provided that the government does not make too big a mess of the areas under its control. What is required for growth is capital, education and technology.
Capital can be accumulated through savings and openness to world capital markets, while technology can be allowed to flow in through foreign direct investment and intellectual openness to the rest of the world. If education, stability and peace can be provided, growth and development should take care of themselves and countries should converge to the level of income that can be supported by the evolution of global technologies. It is this underlying assumption that makes the widening income gaps across countries so puzzling.
A fundamental problem with the standard economic description of the world is that it tries to account for growth and development as a very low dimensional process in which few elements increase in quantity, such as output and physical and human capital.
Output is just GDP, abstracting from the myriad of specific goods and services that are produced. Human capital is just years of schooling, abstracting from the millions of different tasks that individuals and organizations need to master. Physical capital is just a dollar amount, abstracting from the specificity of machines, buildings, power sources, transportation networks and other inputs. Governments provide peace, infrastructure and property rights, abstracting from the millions of pages of legislation they write, edit and extend and the thousands of government agencies that they fund and give marching orders to.
In reality, these aspects of the econosphere are in constant co-evolution and it is the rising complexity of the system that leads to growth and development. Countries evolve by moving to what I call the Adjacent Possible in this complex web of goods and services and productive capabilities. More complex countries have many ways in which they can recombine new capabilities with existing capabilities and products to make new goods and services. Poor countries are trapped by a lack of complexity and limited possibilities of evolving out of it.
Consider the following example: The invention of the computer in 1943 led, some 30 years later, to the opportunity to invent and successfully market the personal computer. In turn, wide sale of the personal computer created jobs, wealth, and opened the opportunity to invent word processing. In turn, word processing offered the opportunity to store word files, which offered the opportunity to share files among scientists at CERN. This technological progress led to the opportunity to invent the World Wide Web. Given the Web, it became possible to use this new niche to market products and eBay flourished. Content assembled on the web, offering the economic opportunity, the new niche, to make profit by inventing search engines and Google has done nicely. Now we have reached the summit of first world civilization with Facebook.
The above examples demonstrate what we know, but have little theory for: economic goods and services and production capacities engender novel goods, services and production capacities in what might be called the Adjacent Possible of the econosphere. We cannot pre-state the way the economy will create new goods and production capacities in the future. Not only do we not know what will happen, we do not even know what can happen.
What is needed is to develop a modified body of theory of economic growth and its policy implications for practical, on the ground, economic experimentation at one or several local or regional economies around the world in the next three years. We must aim at new ideas, not limited by a failed Washington Consensus, while including that which is wise in the standard view.
Are there possible principles governing the emergence of such possibilities? Some mathematical models by myself suggest that an economy with few goods and few production capacities is sub-critical and cannot generate a growing web of new goods and services and production capacities. Ethiopia seems to be sub-critical, producing coffee and thus remaining at the whim of the global coffee market, despite David Ricardo's national competitive advantage in coffee production.
Above a threshold in a coordinate space with diversity of goods on one axis and production capacities on the second, a curved line separates such subcritical economies from supra-critical economies, such as the United States, the European Community, and global economy, that can generate ever novel niches and goods and production capacities. Hausmann and colleagues have recently shown that, in fact, the diversity and richness of the "economic web" in a country correlates with wealth and growth and that, contrary to what emerges from Ricardo's ideas of specialization through comparative advantage, countries diversify as they grow and do so by moving towards the Adjacent Possible as measured by the revealed similarity of products.
Married to the above, standard growth theory, while considering research as a costly and potentially profitable activity which will make new goods, ignores the fact, emphasized by Brian Arthur in The Nature of Technology, that all technologies grow out of existing technologies. Research also exhibits the notion of expanding to its Adjacent Possible. This is the engine that is either trapped in sub-critical economies like Ethiopia, or expanding exponentially in supra-critical economies like China, Korea or the global economy. These are the ideas that are currently missing from the actual design and implementation of growth and development strategies.
We must focus on expanding growth theory to include what we think is missing. Beyond that stated above, there are growing grounds to believe that the way the economy grows cannot be finitely pre-stated. That means that we do not know beforehand the goods and services that may emerge. This is both true, see Silicon Valley and the story leading to Facebook, and it has two major policy implications:
1) Generative environments are needed, and we do not know how to create a supra-critical (local or regional) economy. Clearly, much of development is not related to the expansion of technological possibilities at the global level but the move towards goods and services that are known to the world but previously not yet feasible in a particular country.
2) Because we cannot know how growth will occur, we do not know how the millions of pages of legislation and the thousands of public entities that governments have need to be adjusted to facilitate and accompany the process of change. Standard economic and governmental policy planning, currently conceived as an exercise in adopting unconditional "best practice", need to be revised in ways that are still unclear. Probably, the solution involves thinking of the meta-structure whereby policies co-evolve with capabilities and production.
Thinking of environments in which this process can occur will be a major challenge. The interaction of economic agents, society in general and government is required to reveal the fine-grained information that an effective co-evolutionary process requires. We believe that a new union of standard growth theory, wisdom from the recent development experience, the ideas above, and yet further concepts, can transform our practice of aiding economic growth. The aim should be to have growth experiments on the ground in three years.
Enlarge Romeo Gacad/AFP/Getty Images
Do we need to go beyond the "Washington Consensus" to spur global economic growth?
Romeo Gacad/AFP/Getty Images
Do we need to go beyond the "Washington Consensus" to spur global economic growth?
A body of economic theory known as the "Washington Consensus" has guided attempts to spur global economic growth for over two decades. This Consensus has largely failed, and contemporary economic growth theory seems mostly at a loss to understand the failure.
Meanwhile, the disparity of income between rich and poor countries has grown from 4-to-1 in Adam Smith's time to 70-to-1 now. No one knows why.
Part of the Washington Consensus is based on the work of economist David Ricardo and his theory of national competitive advantage. Ethiopia is good at producing coffee, Alberta is good at producing wheat. The two should trade freely to their joint advantage. Unfortunately, this leaves both as what we call "sub-critical" economies that cannot endogenously generate a growing diversity of goods and production capacities, hence wealth and growth.
We propose an extension of growth theory based on the structure of economic webs and their sub-critical versus supra-critical expansion into an un-prestatable "Adjacent Possible" of the "econosphere" that may be of significant help in guiding attempts to to drive growth.
Four of us, myself, Naresh Singh, acting vice president of the Canadian Partnership Branch of the Canadian International Development Agency, Rohinton Medhora, vice president of programs at the International Development Research Centre in Canada and Ricardo Hausmann, of the Harvard Kennedy School and a past Finance Minister for Venezuela, have co-authored this post and hope, with others, to develop new economic growth theory beyond the Washington Consensus to guide practical economic aid on the ground.
Growth and development strategies are recognized to be in disarray by an increasing number of institutions and stakeholders. It used to be believed that short to-do lists were enough to guide countries to prosperity, whether they be the peace, easy taxes and a tolerable administration of justice of Adam Smith or the openness, sound money and contract enforcement of Larry Summers.
In 1990, John Williamson christened the "Washington Consensus." It advocated fiscal and monetary soundness, openness to trade and investment, financial liberalization and regulation, privatization, deregulation and secure property rights. It was boiled down to a 10-item policy checklist for governments to follow.
After two decades of major reforms in many countries that followed these principles — with generous support from multilateral and bilateral agencies — the results are, in general, surprising and disappointing. The star reformers have not been the star performers and no evidence has been found that these policies promoted growth, once extremely bad policy outcomes are excluded from the sample.
Explaining the disparity in wealth across countries was the question that led Adam Smith to write The Wealth of Nations in 1776. But at that time the income differences to be explained were of the order of 4-to-1. In the meantime, they have grown to over 70-to-1, despite knowledge of the to-do lists and billions spent on aid.
Much of the conventional approach to development strategies is based on the idea that growth and development happen naturally provided that the government does not make too big a mess of the areas under its control. What is required for growth is capital, education and technology.
Capital can be accumulated through savings and openness to world capital markets, while technology can be allowed to flow in through foreign direct investment and intellectual openness to the rest of the world. If education, stability and peace can be provided, growth and development should take care of themselves and countries should converge to the level of income that can be supported by the evolution of global technologies. It is this underlying assumption that makes the widening income gaps across countries so puzzling.
A fundamental problem with the standard economic description of the world is that it tries to account for growth and development as a very low dimensional process in which few elements increase in quantity, such as output and physical and human capital.
Output is just GDP, abstracting from the myriad of specific goods and services that are produced. Human capital is just years of schooling, abstracting from the millions of different tasks that individuals and organizations need to master. Physical capital is just a dollar amount, abstracting from the specificity of machines, buildings, power sources, transportation networks and other inputs. Governments provide peace, infrastructure and property rights, abstracting from the millions of pages of legislation they write, edit and extend and the thousands of government agencies that they fund and give marching orders to.
In reality, these aspects of the econosphere are in constant co-evolution and it is the rising complexity of the system that leads to growth and development. Countries evolve by moving to what I call the Adjacent Possible in this complex web of goods and services and productive capabilities. More complex countries have many ways in which they can recombine new capabilities with existing capabilities and products to make new goods and services. Poor countries are trapped by a lack of complexity and limited possibilities of evolving out of it.
Consider the following example: The invention of the computer in 1943 led, some 30 years later, to the opportunity to invent and successfully market the personal computer. In turn, wide sale of the personal computer created jobs, wealth, and opened the opportunity to invent word processing. In turn, word processing offered the opportunity to store word files, which offered the opportunity to share files among scientists at CERN. This technological progress led to the opportunity to invent the World Wide Web. Given the Web, it became possible to use this new niche to market products and eBay flourished. Content assembled on the web, offering the economic opportunity, the new niche, to make profit by inventing search engines and Google has done nicely. Now we have reached the summit of first world civilization with Facebook.
The above examples demonstrate what we know, but have little theory for: economic goods and services and production capacities engender novel goods, services and production capacities in what might be called the Adjacent Possible of the econosphere. We cannot pre-state the way the economy will create new goods and production capacities in the future. Not only do we not know what will happen, we do not even know what can happen.
What is needed is to develop a modified body of theory of economic growth and its policy implications for practical, on the ground, economic experimentation at one or several local or regional economies around the world in the next three years. We must aim at new ideas, not limited by a failed Washington Consensus, while including that which is wise in the standard view.
Are there possible principles governing the emergence of such possibilities? Some mathematical models by myself suggest that an economy with few goods and few production capacities is sub-critical and cannot generate a growing web of new goods and services and production capacities. Ethiopia seems to be sub-critical, producing coffee and thus remaining at the whim of the global coffee market, despite David Ricardo's national competitive advantage in coffee production.
Above a threshold in a coordinate space with diversity of goods on one axis and production capacities on the second, a curved line separates such subcritical economies from supra-critical economies, such as the United States, the European Community, and global economy, that can generate ever novel niches and goods and production capacities. Hausmann and colleagues have recently shown that, in fact, the diversity and richness of the "economic web" in a country correlates with wealth and growth and that, contrary to what emerges from Ricardo's ideas of specialization through comparative advantage, countries diversify as they grow and do so by moving towards the Adjacent Possible as measured by the revealed similarity of products.
Married to the above, standard growth theory, while considering research as a costly and potentially profitable activity which will make new goods, ignores the fact, emphasized by Brian Arthur in The Nature of Technology, that all technologies grow out of existing technologies. Research also exhibits the notion of expanding to its Adjacent Possible. This is the engine that is either trapped in sub-critical economies like Ethiopia, or expanding exponentially in supra-critical economies like China, Korea or the global economy. These are the ideas that are currently missing from the actual design and implementation of growth and development strategies.
We must focus on expanding growth theory to include what we think is missing. Beyond that stated above, there are growing grounds to believe that the way the economy grows cannot be finitely pre-stated. That means that we do not know beforehand the goods and services that may emerge. This is both true, see Silicon Valley and the story leading to Facebook, and it has two major policy implications:
1) Generative environments are needed, and we do not know how to create a supra-critical (local or regional) economy. Clearly, much of development is not related to the expansion of technological possibilities at the global level but the move towards goods and services that are known to the world but previously not yet feasible in a particular country.
2) Because we cannot know how growth will occur, we do not know how the millions of pages of legislation and the thousands of public entities that governments have need to be adjusted to facilitate and accompany the process of change. Standard economic and governmental policy planning, currently conceived as an exercise in adopting unconditional "best practice", need to be revised in ways that are still unclear. Probably, the solution involves thinking of the meta-structure whereby policies co-evolve with capabilities and production.
Thinking of environments in which this process can occur will be a major challenge. The interaction of economic agents, society in general and government is required to reveal the fine-grained information that an effective co-evolutionary process requires. We believe that a new union of standard growth theory, wisdom from the recent development experience, the ideas above, and yet further concepts, can transform our practice of aiding economic growth. The aim should be to have growth experiments on the ground in three years.
среда, 22 сентября 2010 г.
ИГРЫ СЛОЖНОСТИ
Георгий Малинецкий, Алексей Потапов
В тысячах ситуаций, где самые современные алгоритмы и суперкомпьютеры пасуют, интуиция и опыт позволяют найти разумный компромисс. Почему? Потому что мозг обладает поразительной способностью упрощать мир, выбирать ключевые переменные, самые главные процессы и причинно-следственные связи, верную проекцию реальности. Причем в разных ситуациях разную! Развитие вычислительных систем показало, что эта способность гораздо более удивительна, чем загадочная архитектура мозга, позволяющего решать нестандартные задачи, или таинственная память, преподносящая парадоксальные ассоциации
Если мозг выбирает самое главное и существенное из огромного пространства нашей реальности, значит, оно в ней есть.
Однако, если мозг выбирает самое главное и существенное из огромного фазового пространства нашей реальности, значит, оно в ней есть. Но тогда это можно отобразить в математических моделях. (Если Декарт говорил, что он мыслит, следовательно, существует, модельер может сказать, что он понимает, если может построить математическую модель.) Однако классические математические модели не приспособлены к резким изменениям ситуации — все, что может случиться, фактически уже заложено в модель при ее создании (поэтому модели и позволяют делать открытия). Однако создавать слишком сложные модели, которые содержали бы сразу же все, бессмысленно. Как показывает опыт математического моделирования, их невозможно будет проанализировать. Где же выход?
Итак, в нашем фазовом пространстве есть джокеры. От них одни стараются держаться подальше (как в пословице: «Умный найдет выход из любой ситуации, а мудрый в нее просто не попадет»), а другие активно использовать (знаменитое наполеоновское: «Главное ввязаться в драку, а там посмотрим»). В них неопределенность резко возрастает, а возможности предсказывать дальнейшее уменьшаются. Следовательно, должны быть и другие области, в которых многое или хотя бы самое существенное можно предсказать. Возможно, умение их быстро и точно находить и является главным козырем нашей нервной системы.
Такие области мы будем называть руслами. Название ясно из картинки 7. Близкие траектории как бы притягиваются к некоторому пучку, трубке и далее следуют вместе. Значит, зная детально одну траекторию, можно многое сказать и о других. Политологи, консультанты, референты со времен Римской империи знают, что если в провинции Анчурии заговорили о возрождении национального языка и культуры и о славной истории анчурийского народа, то центр ослаб и большие беспорядки не за горами.
Важно отметить, что картина сближающихся траекторий может наблюдаться не для всех переменных, характеризующих систему, а только для нескольких. Отбрасывая остальные как несущественные («стирая случайные черты»), мы получаем проекцию реальности, в которой ситуация становится предсказуемой, хотя, возможно, с ограниченной точностью и в течение ограниченного промежутка времени. Насколько успешной окажется такая проекция — зависит от системы. Это определяется тем, насколько отброшенное способно повлиять на избранные кандидатуры существенных переменных.
По-видимому, большинство успешных научных теорий приводит к успеху, когда проекция реальности, с которой они имеют дело, оказывается связана с каким-либо руслом (или, если хотите, создание такой успешно предсказывающей теории и показывает, что русло существует и найдено). В идеальном случае очень устойчивых причинно-следственных связей можно оставаться в рамках логики, конструкций типа «Если… то» и навсегда забыть о несущественных деталях. Это и будет обычная математическая модель. Тут раздолье для идеализации, для людей, которые умеют доказывать теоремы.
Джокеры бывают разными. Описание их действий вы найдете в тексте статьи
На следующем уровне находится физика. Ей посчастливилось — она в большинстве ситуаций имеет дело с глобальными руслами, когда можно выделить и описывать почти изолированную подсистему (об окружении можно забыть почти всегда), когда существенными оказываются одни и те же переменные и можно всегда пользоваться одними и теми же уравнениями (дополнительность является скорее исключением, а не правилом). Правда, до теорем обычно дело не доходит, да и на бумажке можно посчитать немного, приходится часто обращаться к помощи компьютера.
В экономике, социологии, психологии, истории ситуация сложнее. Успехи выдающихся экономических теорий, различных психологических школ показывают, что русла есть и здесь. Однако, во-первых, они локальны, то есть обладают предсказывающей силой только в какой-то вполне определенной ситуации. А во-вторых, от них нельзя требовать очень точных и очень длительных прогнозов (хотя от создателей можно потребовать эту точность оценить). Поэтому нужно очень точно оговаривать допущения, исходные посылки. На первый план выходит определение истока (когда посылки начинают быть справедливы) и устья русла (когда они больше не выполняются), определение джокеров — если нельзя указать следующее русло.
Почему русла важны? Потому что понимание, на основе которого можно принимать решения, дают только простые модели, а втиснуть в них действительность можно только отбрасывая «лишнее». По-видимому, на подсознательном уровне мозг решает подобные задачи очень быстро, однако сознательный выбор нужных переменных и его обоснование требует времени, иногда очень большого. Ведь умели же люди очень точно кидать камни и пускать стрелы задолго до Галилея и Ньютона. Мозг быстро прикидывает нужные траекторию и усилия, но никто точно не знает, каким образом. Надо только немного потренироваться. Преуспевающие бизнесмены хорошо ориентируются на биржах и рынках, но обычно не создают экономических теорий и, видимо, почти не пользуются ими. То же самое наблюдается в управлении коллективами людей, сложными объектами, в нетрадиционной медицине и тому подобное. Однако такое эмпирическое знание хотя и приводит к успеху, обычно не может быть передано другим людям, не становится достоянием общества. Его можно передавать только небольшой группе близких соратников личным примером, да и то не всегда. Ученые же теории создают, но социальные теории обычно успешнее всего объясняют прошлое. Пока теория создается, ситуация успевает измениться и старая проекция уже может не отражать сути дела. Найденное русло оказалось пройдено, и текущая ситуация соответствует джокеру или пока не найденному руслу в неизвестной проекции.
Что же делать в такой ситуации? Принципиальным становится определение структуры нашего незнания, осмысление ситуаций, где еще могут существовать русла, и также техники, позволяющей переходить от одних русел к другим, от одних теорий к их альтернативам. Может, к примеру, оказаться, что неокейнсианство и монетаризм — это не альтернативное описание одной реальности, а теории, относящиеся к разным руслам. Поэтому может оказаться, что вопросы «Вы за или…», «Кто прав?» лишены смысла. Следует просто осознать, к какой теории ближе реальность, которую предполагают моделировать или тем более менять. Это необходимо, чтобы не пришлось «импровизировать» или, хуже того, «подгонять» существующую реальность под неадекватную теорию.
Вероятно, именно здесь и может быть развита новая парадигма нелинейной динамики и математического моделирования. Зачем вообще она нужна? Дело в том, что класс объектов, для которых удается строить эффективные модели «из первых принципов», на наш взгляд, в настоящее время почти исчерпан. Для решения многих важных и актуальных задач необходимо строить предсказывающую модель исходя из известной предыстории объекта (подобно тому, как мозг обучается довольно точно бросать камень по результатам тренировочных попыток). И здесь мы встречаемся с серьезнейшими ограничениями на сложность модели. Как показывают некоторые результаты нелинейной динамики, число N здесь обычно не может превышать 5-10. Возможно, следует отказаться (хотя бы частично) от построения общих теорий, а «сосредоточиться лишь на частичном объяснении динамики», на создании «частных теорий», для построения которых было бы достаточно сравнительно небольшого объема информации.
Такого, который может быть собран, переработан и осмыслен за разумный промежуток времени. И для этого концепция русел и джокеров представляется многообещающей. (Здесь уместно будет заметить, что авторы не претендуют на то, что они изобрели нечто принципиально новое. Скорее всего, элементы такого взгляда на научное познание можно найти еще у древних авторов. Мы только хотим подчеркнуть, что предлагаемая концепция позволяет предложить разумное решение ряда серьезных проблем. Именно такой смысл мы вкладываем в слова «третья парадигма».)
Определение русел и джокеров в социальных науках, экологии, теории риска представляется захватывающей задачей. Организация общества, устойчивость и безопасность развития, благополучный внутренний мир выходят на первый план, оттесняя на второй гонку технологий, императивы общества потребления.
Более того, здесь нужен иной уровень междисциплинарного сотрудничества. К сожалению, авторам не раз доводилось сотрудничать на других уровнях. Одни гуманитарии хотели научиться писать украшенные формулами статьи. Другие хотели сначала обсудить методологические проблемы и проверить, можно ли пускать математиков в святая святых. Впрочем, и некоторые наши коллеги-естественники были склонны объяснять, что «многое» в истории, начиная с датировки и кончая никудышной статистикой, следует выбросить на свалку.
Здесь придется учиться слушать и понимать друг друга, искать русла, параметры порядка, проекции реальности.
Речь идет о явлении, которое названо в статье «руслами». Близкие траектории как бы притягиваются к некоторому пучку и далее следуют вместе.
Георгий Малинецкий, Алексей Потапов
В тысячах ситуаций, где самые современные алгоритмы и суперкомпьютеры пасуют, интуиция и опыт позволяют найти разумный компромисс. Почему? Потому что мозг обладает поразительной способностью упрощать мир, выбирать ключевые переменные, самые главные процессы и причинно-следственные связи, верную проекцию реальности. Причем в разных ситуациях разную! Развитие вычислительных систем показало, что эта способность гораздо более удивительна, чем загадочная архитектура мозга, позволяющего решать нестандартные задачи, или таинственная память, преподносящая парадоксальные ассоциации
Если мозг выбирает самое главное и существенное из огромного пространства нашей реальности, значит, оно в ней есть.
Однако, если мозг выбирает самое главное и существенное из огромного фазового пространства нашей реальности, значит, оно в ней есть. Но тогда это можно отобразить в математических моделях. (Если Декарт говорил, что он мыслит, следовательно, существует, модельер может сказать, что он понимает, если может построить математическую модель.) Однако классические математические модели не приспособлены к резким изменениям ситуации — все, что может случиться, фактически уже заложено в модель при ее создании (поэтому модели и позволяют делать открытия). Однако создавать слишком сложные модели, которые содержали бы сразу же все, бессмысленно. Как показывает опыт математического моделирования, их невозможно будет проанализировать. Где же выход?
Итак, в нашем фазовом пространстве есть джокеры. От них одни стараются держаться подальше (как в пословице: «Умный найдет выход из любой ситуации, а мудрый в нее просто не попадет»), а другие активно использовать (знаменитое наполеоновское: «Главное ввязаться в драку, а там посмотрим»). В них неопределенность резко возрастает, а возможности предсказывать дальнейшее уменьшаются. Следовательно, должны быть и другие области, в которых многое или хотя бы самое существенное можно предсказать. Возможно, умение их быстро и точно находить и является главным козырем нашей нервной системы.
Такие области мы будем называть руслами. Название ясно из картинки 7. Близкие траектории как бы притягиваются к некоторому пучку, трубке и далее следуют вместе. Значит, зная детально одну траекторию, можно многое сказать и о других. Политологи, консультанты, референты со времен Римской империи знают, что если в провинции Анчурии заговорили о возрождении национального языка и культуры и о славной истории анчурийского народа, то центр ослаб и большие беспорядки не за горами.
Важно отметить, что картина сближающихся траекторий может наблюдаться не для всех переменных, характеризующих систему, а только для нескольких. Отбрасывая остальные как несущественные («стирая случайные черты»), мы получаем проекцию реальности, в которой ситуация становится предсказуемой, хотя, возможно, с ограниченной точностью и в течение ограниченного промежутка времени. Насколько успешной окажется такая проекция — зависит от системы. Это определяется тем, насколько отброшенное способно повлиять на избранные кандидатуры существенных переменных.
По-видимому, большинство успешных научных теорий приводит к успеху, когда проекция реальности, с которой они имеют дело, оказывается связана с каким-либо руслом (или, если хотите, создание такой успешно предсказывающей теории и показывает, что русло существует и найдено). В идеальном случае очень устойчивых причинно-следственных связей можно оставаться в рамках логики, конструкций типа «Если… то» и навсегда забыть о несущественных деталях. Это и будет обычная математическая модель. Тут раздолье для идеализации, для людей, которые умеют доказывать теоремы.
Джокеры бывают разными. Описание их действий вы найдете в тексте статьи
На следующем уровне находится физика. Ей посчастливилось — она в большинстве ситуаций имеет дело с глобальными руслами, когда можно выделить и описывать почти изолированную подсистему (об окружении можно забыть почти всегда), когда существенными оказываются одни и те же переменные и можно всегда пользоваться одними и теми же уравнениями (дополнительность является скорее исключением, а не правилом). Правда, до теорем обычно дело не доходит, да и на бумажке можно посчитать немного, приходится часто обращаться к помощи компьютера.
В экономике, социологии, психологии, истории ситуация сложнее. Успехи выдающихся экономических теорий, различных психологических школ показывают, что русла есть и здесь. Однако, во-первых, они локальны, то есть обладают предсказывающей силой только в какой-то вполне определенной ситуации. А во-вторых, от них нельзя требовать очень точных и очень длительных прогнозов (хотя от создателей можно потребовать эту точность оценить). Поэтому нужно очень точно оговаривать допущения, исходные посылки. На первый план выходит определение истока (когда посылки начинают быть справедливы) и устья русла (когда они больше не выполняются), определение джокеров — если нельзя указать следующее русло.
Почему русла важны? Потому что понимание, на основе которого можно принимать решения, дают только простые модели, а втиснуть в них действительность можно только отбрасывая «лишнее». По-видимому, на подсознательном уровне мозг решает подобные задачи очень быстро, однако сознательный выбор нужных переменных и его обоснование требует времени, иногда очень большого. Ведь умели же люди очень точно кидать камни и пускать стрелы задолго до Галилея и Ньютона. Мозг быстро прикидывает нужные траекторию и усилия, но никто точно не знает, каким образом. Надо только немного потренироваться. Преуспевающие бизнесмены хорошо ориентируются на биржах и рынках, но обычно не создают экономических теорий и, видимо, почти не пользуются ими. То же самое наблюдается в управлении коллективами людей, сложными объектами, в нетрадиционной медицине и тому подобное. Однако такое эмпирическое знание хотя и приводит к успеху, обычно не может быть передано другим людям, не становится достоянием общества. Его можно передавать только небольшой группе близких соратников личным примером, да и то не всегда. Ученые же теории создают, но социальные теории обычно успешнее всего объясняют прошлое. Пока теория создается, ситуация успевает измениться и старая проекция уже может не отражать сути дела. Найденное русло оказалось пройдено, и текущая ситуация соответствует джокеру или пока не найденному руслу в неизвестной проекции.
Что же делать в такой ситуации? Принципиальным становится определение структуры нашего незнания, осмысление ситуаций, где еще могут существовать русла, и также техники, позволяющей переходить от одних русел к другим, от одних теорий к их альтернативам. Может, к примеру, оказаться, что неокейнсианство и монетаризм — это не альтернативное описание одной реальности, а теории, относящиеся к разным руслам. Поэтому может оказаться, что вопросы «Вы за или…», «Кто прав?» лишены смысла. Следует просто осознать, к какой теории ближе реальность, которую предполагают моделировать или тем более менять. Это необходимо, чтобы не пришлось «импровизировать» или, хуже того, «подгонять» существующую реальность под неадекватную теорию.
Вероятно, именно здесь и может быть развита новая парадигма нелинейной динамики и математического моделирования. Зачем вообще она нужна? Дело в том, что класс объектов, для которых удается строить эффективные модели «из первых принципов», на наш взгляд, в настоящее время почти исчерпан. Для решения многих важных и актуальных задач необходимо строить предсказывающую модель исходя из известной предыстории объекта (подобно тому, как мозг обучается довольно точно бросать камень по результатам тренировочных попыток). И здесь мы встречаемся с серьезнейшими ограничениями на сложность модели. Как показывают некоторые результаты нелинейной динамики, число N здесь обычно не может превышать 5-10. Возможно, следует отказаться (хотя бы частично) от построения общих теорий, а «сосредоточиться лишь на частичном объяснении динамики», на создании «частных теорий», для построения которых было бы достаточно сравнительно небольшого объема информации.
Такого, который может быть собран, переработан и осмыслен за разумный промежуток времени. И для этого концепция русел и джокеров представляется многообещающей. (Здесь уместно будет заметить, что авторы не претендуют на то, что они изобрели нечто принципиально новое. Скорее всего, элементы такого взгляда на научное познание можно найти еще у древних авторов. Мы только хотим подчеркнуть, что предлагаемая концепция позволяет предложить разумное решение ряда серьезных проблем. Именно такой смысл мы вкладываем в слова «третья парадигма».)
Определение русел и джокеров в социальных науках, экологии, теории риска представляется захватывающей задачей. Организация общества, устойчивость и безопасность развития, благополучный внутренний мир выходят на первый план, оттесняя на второй гонку технологий, императивы общества потребления.
Более того, здесь нужен иной уровень междисциплинарного сотрудничества. К сожалению, авторам не раз доводилось сотрудничать на других уровнях. Одни гуманитарии хотели научиться писать украшенные формулами статьи. Другие хотели сначала обсудить методологические проблемы и проверить, можно ли пускать математиков в святая святых. Впрочем, и некоторые наши коллеги-естественники были склонны объяснять, что «многое» в истории, начиная с датировки и кончая никудышной статистикой, следует выбросить на свалку.
Здесь придется учиться слушать и понимать друг друга, искать русла, параметры порядка, проекции реальности.
Речь идет о явлении, которое названо в статье «руслами». Близкие траектории как бы притягиваются к некоторому пучку и далее следуют вместе.
Подписаться на:
Сообщения (Atom)
