Artificial intelligence in the lab: ask not what your computer can do for you
Dick de Ridder
- Year
- 2018
- Citations
- 11
- Access
- Open access
Abstract
In 1957, Herbert Simon, a pioneer of artificial intelligence, predicted that a computer would be the world chess champion within 10 years. It took somewhat longer, but he was eventually proven right when IBM's Deep Blue computer beat Gary Kasparov in 1997. This major breakthrough in artificial intelligence was, in a way, also one of the last successes of what was known as ‘good old-fashioned AI’: the idea that to mimic and understand human intelligence, computers should represent knowledge as symbols and apply reasoning and rules to infer new knowledge. This notion had been criticized for some time already (Dreyfus and Dreyfus, 1992) and over the years gradually lost ground to another approach, machine learning, in which statistical models were fitted to data to derive patterns and correlations. Widely known systems that fit this category include Watson, which successfully competed against the best human players in the Jeopardy! general knowledge quiz, and Google's AlphaGo, which in 2017 beat the reigning world champion at the game of Go. In other settings as well, machine learning progressed. In 2012, it was demonstrated how an extremely large neural network, AlexNet, could be trained to recognize images in 1000 different categories, an atpproach that became known as deep learning (LeCun et al., 2015). Machine learning and deep learning are now routinely used by companies such as Google, Facebook, Amazon and Tesla in products ranging from automated translation and home automation to self-driving cars. In biology, machine learning has likewise found its use. Large volumes of -omics data can now routinely be measured and are used to infer biological function. Many bioinformatics algorithms under the hood rely on statistical models trained on such data to predict – often from nucleotide or amino acid sequences – the structure of genes, the function, location, domain content and secondary structure of proteins, the interactions of proteins with other proteins and DNA, phenotypes, etc. Deep learning has been applied to biological data as well, predicting, among others, protein–DNA interactions (DeepBind), gene regulation (DeepChrome) and variant effects (DeepSEA; Min et al., 2017). A particularly interesting application of machine learning is one where the computer is not only able to predict the function of a sequence or set of sequences, but to (re)design sequences to achieve a certain desired function. This will make it possible to design bespoke regulatory elements, molecules and interactions on-demand, the building blocks needed to fulfil the promise of synthetic biology to engineer microbial machines (Vickers, 2017). Algorithms have been developed to directly predict a sequence given a function, for example inferring an amino acid sequence most likely to fold into a desired three-dimensional structure (O'Connell et al., 2018) or a DNA sequence most likely to bind a certain protein (Killoran et al., 2017). Alternatively, search algorithms can iteratively try mutating a given sequence and keep those changes considered beneficial by a function predictor (Guimaraes et al., 2014), for example to improve protein production (van den Berg et al., 2014). In essence, such sequence (re)design approaches are similar to the AlphaGo setup, in which deep learning networks are used to evaluate Go board positions and moves, based on which a search algorithm decides the next best move to make. In both cases, the search space (the number of mutations or moves to consider) is extremely high-dimensional: in the order of 20400 for a 400-amino acid protein design, and 250150 for a game Go. Deep learning can still successfully learn to predict the value of previously unseen input in such huge spaces, but it requires two things: massive computational resources and extremely large sets of examples. Both are now available for many applications; deep learning was in large part made possible by the advent of affordable GPU-based devices and is most succ
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002