As We May Think - Wikipedia
Bush's 1945 memex sketched hypertext, associative trails, and the personal knowledge machine decades before the hardware to build one existed.
Publication date
Issue
Bush's 1945 memex sketched hypertext, associative trails, and the personal knowledge machine decades before the hardware to build one existed.
Publication date
Sutton's blunt case that general methods riding raw compute keep beating hand-crafted human knowledge, the essay every scaling argument eventually cites.
In computer chess, the methods that defeated the world champion, Kasparov, in 1997, were based on massive, deep search. At the time, this was looked upon with dismay by the majority of computer-chess researchers who had pursued methods that leveraged human understanding of the special structure of chess. When a simpler, search-based approach with special hardware and software proved vastly more effective, these human-knowledge-based chess researchers were not good losers. They said that ``brute force" search may have won this time, but it was not a general strategy, and anyway it was not how people played chess. These researchers wanted methods based on human input to win and were disappointed when they did not.
Turing trades 'can machines think?' for a concrete imitation game and pre-answers nine objections we are somehow still raising in 2026.
"Computing Machinery and Intelligence" is a paper written by Alan Turing on the topic of artificial intelligence. The paper, published in 1950 in Mind, was the first to introduce his concept of what is now known as the Turing test to the general public.
The 2017 paper that threw out recurrence for pure attention and became the architecture underneath every large language model since.
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
Brooks' law, adding people to a late project makes it later, and the enduring lesson that software effort does not parallelize like brick-laying.
From Wikipedia, the free encyclopedia
Alan Kay's 1972 spec for a personal computer for children of all ages, the template every tablet and laptop is still quietly chasing.
From Wikipedia, the free encyclopedia
The field keeps rediscovering its founding ideas. This week: the memex, the imitation game, and the bitter arithmetic of scale.