Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 1 of 20 EXHIBIT 4 Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 2 of 20 I 1111111111111111 1111111111111111 IIIIIIIIIIIIII1111111111111111 Illlllll llll US00RE49461E c19) United States c12) Reissued Patent (10) Patent Number: US RE49,461 E Gorchetchnikov et al. (45) Date of Reissued Patent: Mar. 14, 2023 (54) GRAPHIC PROCESSOR BASED (56) References Cited ACCELERATOR SYSTEM AND METHOD U.S. PATENT DOCUMENTS (71) Applicant: Neurala, Inc., Boston, MA (US) 5,063,603 A 11/1991 Burt (72) Inventors: Anatoli Gorchetchnikov, Newton, MA 5,136,687 A 8/1992 Edelman et al. (US); Heather Marie Ames, Milton, (Continued) MA (US); Massimiliano Versace, Milton, MA (US); Fabrizio Santini, FOREIGN PATENT DOCUMENTS Jamaica Plain, MA (US) EP 1224622 Bl 11/2004 WO 190208 11/2014 (73) Assignee: Neurala, Inc., Boston, MA (US) (Continued) (21) Appl. No.: 17/136,343 (22) Filed: Dec. 29, 2020 OTHER PUBLICATIONS Related U.S. Patent Documents Hodgkin, A. L., and Huxley, A. F. 1952. Quantitative description of Reissue of: membrane current and its application to conduction and excitation (64) Patent No.: 9,189,828 m nerve. J Physiol 117, pp. 500-544. Issued: Nov. 17, 2015 (Continued) Appl. No.: 14/147,015 Primary Examiner - William H. Wood Filed: Jan.3, 2014 (74) Attorney, Agent, or Firm - Smith Baluch LLP U.S. Applications: (63) Continuation of application No. 15/808,201, filed on (57) ABSTRACT Nov. 9, 2017, now Pat. No. Re. 48,438, which is an An accelerator system is implemented on an expansion card (Continued) comprising a printed circuit board having (a) one or more graphics processing units (GPUs), (b) two or more associ- (51) Int. Cl. ated memory banks (logically or physically partitioned), (c) G06T 1160 (2006.01) a specialized controller, and (d) a local bus providing signal coupling compatible with the PCI industry standards. The G06F 9/50 (2006.01) controller handles most of the primitive operations to set up (Continued) and control GPU computation. Thus, the computer's central (52) U.S. Cl. processing unit (CPU) can be dedicated to other tasks. In this CPC .............. G06T 1120 (2013.01); G06F 9/5027 case a few controls (simulation start and stop signals from the CPU and the simulation completion signal back to CPU), (2013.01); G06T 1160 (2013.01); GPU programs and input/output data are exchanged between (Continued) CPU and the expansion card. Moreover, since on every time (58) Field of Classification Search step of the simulation the results from the previous time step CPC ... G06F 9/5027; G06F 2209/509; G06T 1/20; are used but not changed, the results are preferably trans- ferred back to CPU in parallel with the computation. G06T 1/60; G06N 3/00; G06N 3/02; (Continued) 21 Claims, 5 Drawing Sheets Expansion Card mo I ............ 41 r:;- r~·--· .tll) Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 3 of 20 US RE49,461 E Page 2 Related U.S. Application Data 2011/0004341 Al 1/2011 Sarvadevabhatla et al. 2011/0173015 Al 7/2011 Chapman et al. application for the reissue of Pat. No. 9,189,828, 2011/0279682 Al 11/2011 Li et al. which is a continuation of application No. 11/860, 2012/0072215 Al 3/2012 Yu et al. 2012/0089552 Al 4/2012 Chang et al. 254, filed on Sep. 24, 2007, now Pat. No. 8,648,867. 2012/0197596 Al 8/2012 Comi (60) Provisional application No. 60/826,892, filed on Sep. 2012/0316786 Al 12/2012 Liu et al. 2013/0126703 Al 5/2013 Caulfield 25, 2006. 2013/0131985 Al 5/2013 Weiland et al. 2014/0019392 Al 1/2014 Buibas et al. (51) Int. Cl. 2014/0032461 Al 1/2014 Weng G06T 1120 (2006.01) 2014/0052679 Al 2/2014 Sinyavskiy et al. G06N 3/063 (2023.01) 2014/0089232 Al 3/2014 Buibas et al. 2015/0127149 Al 5/2015 Sinyavskiy et al. G06N 20/00 (2019.01) 2015/0134232 Al 5/2015 Robinson (52) U.S. Cl. 2015/0224648 Al 8/2015 Lee et al. CPC ........ G06F 2209/509 (2013.01); G06N 3/063 2016/0075017 Al 3/2016 Laurent et al. (2013.01); G06N 20/00 (2019.01) 2016/0082597 Al 3/2016 Gorshechnikov et al. 2016/0096270 Al 4/2016 Gabardos et al. (58) Field of Classification Search 2016/0198000 Al 7/2016 Gorshechnikov et al. CPC ............ G06N 3/04; G06N 3/06; G06N 3/063; 2017 /0024877 Al 1/2017 Versace et al. G06N 3/08; G06N 3/1 O; G06N 5/00; 2017/0076194 Al 3/2017 Versace et al. G06N 7/00; G06N 7/02; G06N 7/04; 2017/0193298 Al 7/2017 Versace et al. G06N 7/046; G06N 7/06; G06N 20/00 See application file for complete search history. FOREIGN PATENT DOCUMENTS WO 2014204615 A2 12/2014 (56) References Cited WO 2015143173 A2 9/2015 WO 2016014137 A2 1/2016 U.S. PATENT DOCUMENTS 5,142,665 A 8/ 1992 Bigus OTHER PUBLICATIONS 5,172,253 A 12/1992 Lynne 5,388,206 A 2/1995 Poulton et al. Hopfield, J. 1982. Neural networks and physical systems with 6,018,696 A 1/2000 Matsuoka et al. emergent collective computational abilities. In Proc Natl Acad Sci 6,336,051 Bl 1/2002 Pangels et al. 6,647,508 B2 11/2003 Zalewski et al. USA, vol. 79, pp. 2554-2558. 7,119,810 B2 10/2006 Sumanaweera et al. Ilie, A. 2002. Optical character recognition on graphics hardware. 7,219,085 B2 * 5/2007 Buck . G06V 10/955 Tech. Rep. integrative paper, UNCCH, Department of Computer 706/12 Science, 9 pages. 7,477,256 Bl* 1/2009 Johnson .................... G06F 3/14 International Preliminary Report on Patentability in related PCT 345/506 Application No. PCT/US2014/039162 filed May 22, 2014, dated 7,525,547 Bl* 4/2009 Diard ........................ G06T 1/20 Nov. 24, 2015, 7 pages. 345/522 7,765,029 B2 7/2010 Fleischer et al. International Preliminary Report on Patentability in related PCT 7,861,060 Bl* 12/2010 Nickolls et al. .. ... ... .... ... . 712/22 Application No. PCT/US2014/039239 filed May 22, 2014, dated 7,873,650 Bl 1/2011 Chapman et al. Nov. 24, 2015, 8 pages. 8,392,346 B2 3/2013 Ueda et al. International Preliminary Report on Patentability dated Nov. 8, 8,510,244 B2 8/2013 Carson et al. 2016 from International Application No. PCT/US2015/029438, 7 8,583,286 B2 11/2013 Fleischer et al. pages. 8,648,867 B2 * 2/2014 Gorchetchnikov et al ... 345/501 International Search Report and Written Opinion dated Feb. 18, 9,031,692 B2 5/2015 Zhu 9,177,246 B2 11/2015 Bui bas et al. 2015 from International Application No. PCT/US2014/039162, 12 9,189,828 B2 11/2015 Gorchetchnikov et al. pages. 9,626,566 B2 4/2017 Versace et al. International Search Report and Written Opinion dated Feb. 23, 10,083,523 B2 9/2018 Versace et al. 2016 from International Application No. PCT/US2015/029438, 11 RE48,438 E * 2/2021 Gorchetchnikov ... G06F 9/5027 pages. 2001/0010034 Al 7/2001 Burton International Search Report and Written Opinion dated Jul. 6, 2017 2002/0046271 Al 4/2002 Huang from International Application No.PCT/US2017/029866, 12 pages. 2002/0050518 Al 5/2002 Roustaei 2002/0064314 Al 5/2002 Comaniciu et al. International Search Report and Written Opinion dated Nov. 26, 2002/0168100 Al 11/2002 Woodall 2014 from International Application No. PCT/US2014/039239, 14 2003/0026588 Al 2/2003 Elder et al. pages. 2003/00787 54 Al 4/2003 Hamza International Search Report and Written Opinion dated Sep. 15, 2004/0015334 Al 1/2004 Ditlow et al. 2015 from International Application No. PCT/US2015/021492, 9 2005/0166042 Al 7/2005 Evans pages. 2006/0129506 Al* 6/2006 Edelman. G05D 1/0088 Itti, L., and Koch, C. (2001). Computational modelling of visual 706/12 2006/0184273 Al 8/2006 Sawada et al. attention. Nature Reviews Neuroscience, 2 (3), 194-203. 2007/0052713 Al 3/2007 Chung et al. Itti, L., Koch, C., and Niebur, E. (1998). A Model of Saliency-Based 2007/0198222 Al 8/2007 Schuster et al. Visual Attention for Rapid Scene Analysis, 1-6. 2007/0279429 Al 12/2007 Ganzer Jarrett, K., Kavukcuoglu, K., Ranzato, M. A., & LeCun, Y. (Sep. 2008/0033897 Al 2/2008 Lloyd 2009). What is the best multi-stage architecture tor object recogni- 2008/0066065 Al 3/2008 Kim et al. tion?. In Computer Vision, 2009 IEEE 12th International Confer- 2008/0258880 Al 10/2008 Smith et al. ence on (pp. 2146-2153) IEEE. 2009/0080695 Al 3/2009 Yang 2009/0089030 Al 4/2009 Sturrock et al. Khaligh-Razavi, S.-M et al., Deep Supervised, but Not Unsuper- 2009/0116688 Al 5/2009 Monacos et al. vised, Models May Explain IT Cortical Representation, PLoS 2010/0048242 Al 2/2010 Rhoads et al. Computational Biology, vol. 10, Issue 11, 29 pages (Nov. 2014). Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 4 of 20 US RE49,461 E Page 3 (56) References Cited Minih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, OTHER PUBLICATIONS Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529- Kim, S., Novel approaches to clustering, biclustering and algo- 533, Feb. 25, 2015. rithms based on adaptive resonance theory and Intelligent control, Montrym et al., The GeForce 6800, in IEEE Micro, vol. 25, No. 2, Doctoral Dissertations, Missouri University of Science and Tech- pp. 41-51, March-Apr. 2005. nology, 125 pages (2016). Moore, Andrew W and Atkeson, Christopher G. Prioritized sweep- Kipfer, P., Segal, M., and Westermann, R. 2004. UberFlow: A ing: Reinforcement learning with less data and less time. Machine GPU-Based Particle Engine. In Proceedings of the SIGGRAPH/ Learning, 13(1): 103-130,1993. Eurographics Workshop on Graphics Hardware 2004, pp. 115-122. Najernnik, J., and Geisler, W. (2009). Simple sununation rule for Kolb, A., L. Latta, and C. RF7K-SALAMA. 2004. "Hardware- optimal fixation selection in visual search. Vision Research. 49, Based Simulation and Collision Detection for Large Particle Sys- 1286-1294. tems." In Proceedings of the SIGGRAPH/Eurographics Workshop Non-Final Office Action dated Jan. 4, 2018 from U.S. Appl. No. on Graphics Hardware 2004, pp. 123-131. 15/262,637, 23 pages. Kompella, Varun Raj, Luciw, Matthew, and Schmidhuber, Jurgen. Non-Final Office Action dated May 31, 2018 from U.S. Appl. No. Incremental slow feature analysis: Adaptive low-complexity slow 14/947,516, 16 pages. feature updating from high-dimensional input streams Neural Com- Notice of Alllowance dated May 22, 2018 from U.S. Appl. No. putation, 24( 11):2994-3024, 2012. 15/262,637, 6 pages. Kowler, E. (2011). Eye movements: The past 25years. Vision Notice of Allowance dated Jul. 27, 2016 from U.S. Appl. No. Research, 51(13), 1457-1483. doi:10.1016/j.visres.2010.12.014. 14/662,657. Larochelle H., & Hinton G. (2012). Learning to combine foveal Notice of Allowance dated Dec. 16, 2016 from U.S. Appl. No. glimpses with a third-order Boltzmann machine. NIPS 2010,1243- 14/662,657. 1251. Oh, K.-S., and Jung, K. 2004. GPU implementation of neural LeCun, Y., Kavukcuoglu, K., & Farabet, C. (May 2010). Convolu- networks. Pattern Recognition 37, pp. 1311-1314. tional networks and applications in vision. In Circuits and Systems Oja, E. (1982). Simplified neuron model as a principal component (ISCAS), Proceedings of 2010 IEEE International Symposium on analyzer. Journal of Mathematical Biology 15(3), 267-273. (pp. 253-256). IEEE. Partial Supplementary European Search Report dated Jul. 4, 2017 Lee, D. D. and Seung, H. S. (1999). Learning the parts of objects from European Application No. 14800348.6, 13 pages. by non-negative matrix factorization. Nature, 401(6755):788-791. Perumalla, "Discrete-event execution alternatives on general pur- Lee, D. D., and Seung, H. S. (1997). "Unsupervised learning by pose graphical processing units (GPGPU s)." Proceedings of the convex and conic coding." Advances in Neural Information Pro- 20th Workshop on Principles of Advanced and Distributed Simu- cessing Systems, 9. lation. IEEE Computer Society, 2006.8 pages. Legenstein, R., Wilbert, N., and Wiskott, L. Reinforcement learning Raijmakers, M.E.J., and Molenaar, P. (1997). Exact Art: A complete on slow features of high-dimensional input streams. PLoS Compu- implementation of an ART network Neural networks 10 (4), 649- tational Biology, 6(8), 2010. ISSN 1553-734X. 13 pages. 669. Leveille, J., Ames, H., Chandler, B., Gorchetchnikov, A., Mingolla, Ranzato, M.A., Huang, F. J., Boureau, Y. L., & Lecun, Y. (2007, E., Patrick, S., and Versace, M. (2010) Learning in a distributed June). Unsupervised learning of invariant feature hierarchies with software architecture for large-scale neural modeling. BIONET- applications to object recognition. In Computer Vision and Pattern ICS 10, Boston, MA, USA. 8 pages. Recognition, 2007. CVPR'07. IEEE Conference on (pp. 1-8). IEEE. Livitz G., Versace M., Gorchetchnikov A., Vasilkoski Z., Ames H., Raudies, F., Eldridge, S., Joshi, A., and Versace, M. (Aug. 20, 2014). Chandler B., Leveille J. andMingolla E. (2011) Adaptive, brain-like Learning to navigate in a virtual world using optic flow and stereo systems give robots complex behaviors, The Neuromorphic Engi- disparity signals. Artificial Life and Robotics, DOI 10.1007/10015- neer,: 10.2417/1201101.003500 Feb. 2011. 3 pages. 014-0153-l. 15 pages. Livitz, G., Versace, M., Gorchetchnikov, A., Vasilkoski, Z., Ames, Adelson, E. H, Anderson, C. H, Bergen, JR., Burt, P. J, & Ogden, H., Chandler, B., Leveille, J., Mingolla, E., Snider, G., Amerson, R., J. M (1984) Pyramid methods in image processing. RCA engineer, Carter, D., Abdalla, H., and Qureshi, S. (2011) Visually-Guided 29(6), 33-41. Adaptive Robot (ViGuAR). Proceedings of the International Joint Aggarwal, Charu C, Hinneburg, Alexander, and Keim, Daniel A. On Conference on Neural Networks (IJCNN) 2011, San Jose, CA, the surprising behavior of distance metrics in high dimensional USA. 9 pages. space. Springer, 2001. 15 pages. Al-Kaysi, A. M. et al., A Multichannel Deep Belief Network for the Lowe, D.G.(2004). Distinctive Image Features from Scale-Invariant Classification of EEG Data, from Ontology-based Information Keypoints. Journal International Journal of Computer Vision archive Extraction for Residential Land Use Suitability: A Case Study of the vol. 60, 2, 91-110. City of Regina, Canada, DOI 10.1007/978-3-319-26561-2_5, 8 Lu, Z.L., Liu, J., and Dosher, B.A.(2010) Modeling mechanisms of pages (Nov. 2015). perceptual learning with augmented Hebbian re-areighting Vision Ames, H, Versace, M., Gorchetchnikov, A., Chandler, B., Livitz, G., Research, 50(4). 375-390. Leveille, J., Mingolla, E., Carter, D., Abdalla, H., and Snider, G. Luo et al., "Ailificial neural network computation on graphic (2012) Persuading computers to act more like brains. In Advances process unit." Neural Networks, 2005. IJCNN'05. Proceedings in Neuromorphic Mernristor Science and Applications, Kozma, 2005 IEEE International Joint Conference on vol. 1 IEEE, 2005 pp. R.Pino,R., and Pazienza, G. (eds), Springer Verlag. 25 pages. 622-626. Ames, H. Mingolla, E., Sohail, A., Chandler, B., Gorchetchnikov, Mahadevan, S. Proto-value functions: Developmental reinforce- A., Leveille, J., Livitz, G. and Versace, M. (2012) The Animat. IEEE ment learning. In Proceedings of the 22nd international conference Pulse, Feb. 2012, 3(1), 47-50. on Machine learning, pp. 553-560. ACM, 2005. Apolloni, B. et al., Training a network of mobile neurons, Proceed- Meuth, J.R. and Wunsch, D.C. (2007) A Survey of Neural Compu- ings of International Joint Conference on Neural Networks, San tation On Graphics Processing Hardware. 22nd IEEE International Jose, CA, doi: 10.1109/IJCNN.2011.6033427, pp. 1683-1691 (Jul. Symposium on Intelligent Control, Part of IEEE Multi-conference 31-Aug. 5, 2011). on Systems and Control, Singapore, Oct. 1-3, 2007, 5 pages. Artificial Intelligence as a Service. Invited talk, Defrag, Broomfield, Mishkin M, Ungerleider LG. (1982). "Contribution of striate inputs CO, Nov. 4-6, 2013. 22 pages. to the visuospatial functions of parieto-preoccipital cortex in mon- Aryananda, L. 2006. Attending to learn and learning to attend for a keys," Behav Brain Res, 6 (1): 57-77. social robot. Humanoids 06, pp. 618-623. Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 5 of 20 US RE49,461 E Page 4 (56) References Cited Ellias, S. A., and Grossberg, S. 1975. Pattern formation, contrast control and oscillations in the short term memory of shunting OTHER PUBLICATIONS on-center off-surround networks Biol Cybern 20, pp. 69-98. Extended European Search Report and Written Opinion dated Jun. Baraldi, A. and Alpaydin, E. ( 1998). Simplified Art: A new class of 1, 2017 from European Application No. 14813864.7, 10 pages. Art algorithms. International Computer Science Institute, Berkeley, Extended European Search Report and Written Opinion dated Oct. CA, TR-98-004, 1998. 42 pages. 12, 2017 from European Application No. 14800348.6, 12 pages. Baraldi, A. and Alpaydin, E. (2002). Constructive feedforward Art Extended European Search Report and Written Opinion Oct. 23, clustering networks-Part I. IEEE Transactions on Neural Net- 2017 from European Application No. 15765396.5, 3 pages. works 13(3), 645-661. Faz!, A., Grossberg, S., and Mingolla, E. (2009). View-invariant Baraldi, A. and Parmiggiani, F. (1997). Fuzzy combination of object category learning, recognition, and search: How spatial and Kohonen's and ART neural network models to detect statistical object attention are coordinated using surface-based attentional regularities in a random sequence of multi-valued input patterns. In shrouds. Cognitive Psychology 58, 1-48. International Conference on Neural Networks, IEEE. 6 pages. Fiildiak, P. (1990). Forming sparse representations by local anti- Baraldi, Andrea and Alpaydin, Ethem. Constructive feedforward Hebbian learning, Biological Cybernetics, vol. 64, pp. 165-170. ART clustering networks-part II IEEE Transactions on Neural Friston K., Adams R., Perrinet L., & Breakspear M. (2012). Per- Networks, 13(3):662-677, May 2002. ISSN 1045-9227. doi: 10.1109/ ceptions as hypotheses: saccades as experiments. Frontiers in Psy- tnn.2002.1000131. URL http://dx.doi.org/l 0.1109/tnn.2002. chology, 3 (151), 1-20. 1000131. Galbraith, B.V, Guenther, F.H., and Versace, M. (2015) A neural Bengio, Y., Courville, A., & Vincent, P. Representation learning: A network-based exploratory learning and motor planning system for review and new perspectives, IEEE Transactions on Pattern Analy- co-robots Frontiers in Neuroscience, in press. 10 pages. sis and Machine Intelligence, vol. 35 Issue 8, Aug. 2013. pp. George, D. and Hawkins, J. (2009). Towards a mathematical theory 1798-1828. of cortical micro-circuits. PLoS Computational Biology 5( 10), 1-26. Berenson, D et al., A robot path planning framework that learns Georgii, J., and Westermann, R. 2005. Mass-spring systems on the from experience, 2012 International Conference on Robotics and GPU. Simulation Modelling Practice and Theory 13, pp. 693-702. Automation, 2012, 9 pages [retrieved from the internet] URL:http:// Gorchetchnikov A., Hasselmo M. E. (2005). A biophysical imple- users. wpi .edu/-dberenson/lightning. pdf. mentation of a bidirectional graph search algorithm to solve mul- Bernhard, F., and Keriven, R. 2005. Spiking Neurons on GPUs. tiple goal navigation tasks. Connection Science, 17(1-2), pp. 145- Tech. Rep. 05-15, Ecole Nationale des Ponts et Chauss'es, 8 pages. 166. Bes!, P. J., & Jain, R. C. (1985). Three-dimensional object recog- Gorchetchnikov A., Hasselmo M. E. (2005). A simple rule for nition. ACM Computing Surveys (CSUR), 17(1), 75-145. spike-timing-dependent plasticity: local influence of AHP current. Boddapati, V., Classifying Environmental Sounds with Image Net- Neurocomputing, 65-66, pp. 885-890. works, Thesis, Faculty of Computing Blekinge Institute of Tech- Gorchetchnikov A., Versace M., Hasselmo M. E. (2005). A Model nology, 37 pages (Feb. 2017). of STDP Based on Spatially and Temporally Local Information: Bohn, C.-A. Kohonen. 1998. Feature Mapping Through Graphics Derivation and Combination with Gated Decay. Neural Networks, Hardware. In Proceedings of 3rd Int. Conference on Computational 18, pp. 458-466. Intelligence and Neurosciences, 4 pages. Gorchetchnikov A., Versace M., Hasselmo M. E. (2005). Spatially Bradski, G., & Grossberg, S. (1995). Fast-learning Viewnet archi- and temporally local spike-timing-dependent plasticity rule. In: tectures for recognizing three-dimensional objects from multiple Proceedings of the International Joint Conference on Neural Net- two-dimensional views. Neural Networks, 8 (7-8), 1053-1080. works, No. 1568 in IEEE CD-ROM Catalog No. 05CH37662C, pp. Canny, J.A. (1986). Computational Approach To Edge Detection, 390-396. IEEE Trans. Pattern Analysis and Machine Intelligence, 8(6):679- Gorcheichnikov, A. 2017. An Approach to a Biologically Realistic 698. Simulation of Natural Memory. Master's thesis, Middle Tennessee Carpenter, G.A. and Grossberg, S. (1987). A massively parallel State University, Murfreesboro, TN, 70 pages. architecture for a self-organizing neural pattern recognition machine. Grossberg, S. (1973). Contour enhancement, short-term memory, Computer Vision, Graphics, and Image Processing 37, 54-115. and constancies in reverberating neural networks. Studies in Applied Carpenter, G.A., and Grossberg, S. (1995). Adaptive resonance Mathematics 52, 213-257. theory (ART). In M. Arbib (Ed.), The handbook of brain theory and Grossberg, S., and Huang, T.R. (2009). Artscene: A neural system neural networks, (pp. 79-82). Cambridge, M.A.: MIT press. for natural scene classification. Journal of Mision, 9 (4), 6.1-19. Carpenter, G.A., Grossberg, S. and Rosen, D.B. (1991). Fuzzy Art: doi: 10.1167/9.4.6. Fast stable learning and categorization of analog patterns by an Grossberg, S., and Versace, M. (2008) Spikes, synchrony, and adaptive resonance system Neural Networks 4, 759-771. attentive learning by laminar thalamocortical circuits. Brain Research, Carpenter, Gail A and Grossberg, Stephen. The art of adaptive 1218C, 278-312 [Authors listed alphabetically]. pattern recognition by a self-organizing neural network. Computer, Hagen, T. R., Hjelmervik, J., Lie, K.-A., Natvig, J., and Ofstad 21(3):77-88, 1988. Henriksen, M. 2005. Visual simulation of shallow-water waves. Coifman, R.R. and Maggioni, M. Diffusion wavelets. Applied and Simulation Modelling Practice and Theory 13, pp. 716-726. Computational Harmonic Analysis, 21(1):53-94, 2006. Hasselt, Hado Van. Double q-learning. In Advances in Neural Coifman, R.R., Lafon, S., Lee, A.B., Maggioni, M., Nadler, B., Information Processing Systems, pp. 2613-2621,2010. Warner, F., and Zucker, S.W. Geometric diffusions as a tool for Hinton, G. E., Osindero, S., and Teh, Y. (2006). A fast learning harmonic analysis and structure definition of data: Diffusion maps. algorithm for deep belief nets. Neural Computation, 18, 1527-1554. Proceedings of the National Academy of Sciences of the United Ren, Y et al., Ensemble Classification and Regression-Recent States of America, 102(21):7426, 2005. 21 pages. Developments, Applications and Future Directions, in IEEE Com- Cornwall et al., Automatically translating a general purpose C++ putational Intelligence Magazine, 10.1109/MCI.2015.2471235, 14 image processing library for GPUs. Proceedings 20th IEEE Inter- pages (2016). national Parallel & Distributed Processing Symposium, 2006, 8 Riesenhuber, M., & Poggio, T. (1999). Hierarchical models of pages. object recognition in cortex. Nature Neuroscience, 2 (11), 1019- Davis, C. E. 2005. Graphic Processing Unit Computation of Neural 1025. Networks. Master's thesis, University of New Mexico, Albuquer- Riesenhuber, M., & Poggio, T. (2000). Models of object recogni- que, NM, 121 pages. tion. Nature neuroscience, 3, 1199-1204. Dosher, B.A., and Lu, Z.L. (2010). Mechanisms of perceptual Rolfes, T. 2004. Artificial Neural Networks on Programmable attention in preening of location. Vision Res., 40(10-12). 1269- Graphics Hardware. In Game Programming Gems 4, A. Kirmse, Ed. 1292. Charles River Media, Hingham, MA, pp. 373-378. Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 6 of 20 US RE49,461 E Page 5 (56) References Cited Snider, Greg, et al. "From synapses to circuitry: Using mernristive memory to explore the electronic brain." IEEE Computer, vol. OTHER PUBLICATIONS 44(2). (2011): 21-28. Spratling, M. W. (2008). Predictive coding as a model of biased Rublee, E., Rabaud, V., Konolige, K., & Bradski, G. (2011). ORB: competition in visual attention. Vision Research, 48(12): 1391-1408. An efficient alternative to SIFT or SURF. In IEEE International Spratling, M. W. (2012). Unsupervised learning of generative and Conference on Computer Vision (ICCV) 2011, 2564-2571. discriminative weights encoding elementary image components in a Ruesch, J et al. 2008. Multimodal Saliency-Based Bottom-Up predictive coding model of cortical function. Neural Computation, Attention a Framework for the Humanoid Robot iCub. 2008 IEEE 24(1):60-103. International Conference on Robotics and Automation, pp. 962-965. Spratling, M. W., De Meyer, K., and Kompass, R. (2009). Unsu- Rumelhart D., Hinton G., and Williams, R. (1986). Learning inter- pervised learning of overlapping image components using divisive nal representations by error propagation. In Parallel distributed input modulation. Computational intelligence and neuroscience. 20 processing: explorations in the microstructure of cognition, vol. 1, pages. MIT Press. 45 pages. Sprekeler, H. On the relation of slow feature analysis and laplacian Rumpf, M. and Strzodka, R. Graphics processor units: New pros- eigenmaps. Neural Computation, pp. 1-16, 2011. pects for parallel computing. In Are Magnus Bruaset and Aslak Sun, Z et al., Recognition of SAR target based on multilayer Tveito, editors, Numerical Solution of Partial Differential Equations auto-encoder and SNN, International Journal of Innovative Com- on Parallel Computers, vol. 51 of Lecture Notes in Computational puting, Information and Control, vol. 9, No. 11, pp. 4331-4341, Science and Engineering, pp. 89-134. Springer, 2005. Nov. 2013. Salakhutdinov, R., & Hinton, G. E. (2009). Deep boltzmann machines. Sutton, R. S., and Barto, A.G. (1998). Reinforcement learning: An In International Conference on Artificial Intelligence and Statistics introduction(vol. 1, No. 1). Cambridge: MIT press. 10 pages. (pp. 448-455). Tong, F., Ze-Nian Li, (1995). Reciprocal-wedge transform for Schaul, Tom, Quan, John, Antonoglou, Ioannis, and Silver, David. space-variant sensing, Pattern Analysis and Machine Intelligence, Prioritized experience replay. arXiv preprint arXiv: 1511.05952, IEEE Transactions on , vol. 17, No. 5, pp. 500-551 doi: 10 Nov. 18, 2015. 21 pages. 1109/34.391393. Schmidhuber, J. (2010). Formal theory of creativity, fun, and Torralba, A., Oliva, A., Castelhano, M.S., Henderson, J.M. (2006). intrinsic motivation (1990-2010). Autonomous Mental Develop- Contextual guidance of eye movements and attention in real-world ment, IEEE Transactions on, 2(3), 230-247. scenes: the role of global features in object search Psychological Schmidhuber, Jurgen. Curious model-building control systems. In Review, 113(4).766-786. Neural Networks, 1991. 1991 IEEE International Joint Conference Van Hasselt, Hado, Guez, Arthur, and Silver, David. Deep rein- on, pp. 1458-1463. IEEE, 1991. forcement learning with double q-learning. arXiv preprint arXiv: Seibert, M., & Waxman, A.M. (1992). Adaptive 3-D Object Rec- 1509.06461, Sep. 22, 2015. 7 pages. ognition from Multiple Views. IEEE Transactions on Pattern Analy- Versace, Brain-inspired computing. Invited keynote address, Bionet- sis and Machine Intelligence, 14 (2), 107-124. ics 2010, Boston, MA, USA. 1 page. Setoain et al., "Parallel hyperspectral image processing on com- Versace, M. (2006) From spikes to interareal synchrony: how modity graphics hardware." Parallel Processing Workshops, 2006. attentive matching and resonance control learning and Information ICPP 2006 Workshops 2006 International Conference on. IEEE, processing by laminar thalamocortical circuits. NSF Science of 2006. 8 pages. Learning Centers PI Meeting, Washington, DC, USA. 1 page. Sherbakov, L. and Versace, M. (2014) Computational principles for Versace, M., (2010) Open-source software for computational neu- an autonomous active vision system. Ph.D., Boston University, roscience: Bridging the gap between models and behavior. In http://search.proquest.com/docview/l 558856407. 194 pages. Horizons in Computer Science Research, vol. 3 43 pages. Sherbakov, L. et al. 2012. CogEye: from active vision to context Versace,M., Ames, H., Leveille, J., Fortenberry,B., and Gorchetchnikov, identification, youtube, retrieved from the Internet an Oct. 10, 2017: A. (2008) KlnNeSS: A modular framework for computational URL://www.youtube.com/watch?v~i5PQk962Blk, 1 page. neuroscience Neuroinforrnatics, 2008 Winter; 6(4):291-309. Epub Sherbakov, L. et al. 2013. CogEye: system diagram module brain Aug. 10, 2008. area function algorithm approx # neurons, retrieved from the Versace, M., and Chandler, B. (2010) MoNeta: A Mind Made from Internet on Oct. 12, 2017: URL://http://www-labsticc.univ-ubs.fr/ Mernristors. IEEE Spectrum, Dec. 2010. 8 pages. --coussy/neucomp2013/index_fichiers/material/posters/NeuCornp2013_ Versace, TEDx Fulbright, Invited talk, Washington DC, Apr. 5, final56x36.pdf, 1 page. 2014. 30 pages. Sherbakov, L., Livitz, G., Sohail, A., Gorchetchnikov, A., Mingolla, Webster, Bachevalier, Ungerleider (1994). Connections of IT areas E., Ames, H., and Versace, M (2013b) A computational model of the TEO and TE with parietal and frontal cortex in macaque monkeys. role of eye-movements in object disambiguation. Cosyne, Feb. Cerebal Cortex, 4(5), 470-483. 28-Mar. 3, 2013. Salt Lake City, UT, USA. 2 pages. Wiskott, Laurenz and Sejnowski, Terrence. Slow feature analysis: Sherbakov, L., Livitz, G., Sohail, A., Gorchetchnikov, A., Mingolla, Unsupervised learning ofinvariances. Neural Computation, 14(4):715- E., Ames, H., and Versace, M. (2013a) CogEye: An online active 770, 2002. vision system that disambiguates and recognizes objects NeuComp Wu, Yan & J. Cai, H. (2010). A Simulation Study of Deep Belief 2013.2 pages. Network Combined with the Self-Organizing Mechanism of Adap- Smolensky, Paul. Information processing in dynamical systems: tive Resonance Theory. 10.1109/CISE.2010.56//265, 4 pages. Foundations of harmony theory. No. CU-CS-321-86. Colorado Univ At Boulder Dept of Computer Science, 1986. 88 pages. * cited by examiner Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 7 of 20 U.S. Patent Mar.14,2023 Sheet 1 of 5 US RE49,461 E ,/ ........... ·····t--·.c: Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 8 of 20 U.S. Patent Mar.14,2023 Sheet 2 of 5 US RE49,461 E w rt: ~ X: \,,, ;- 0 :) 2; ..., X < ~ ~ d w w 00 :t, ..-. -0 0 i,t:< N \{) (~, ..1 r- :fil <.(} t>J ':") ;-b (~ / CJ "<...... N ~~ >- ~ 0 0 N ., <".\! - u !! w ..J V '? / z.U ,.J. 0 :) l. :Ez ~ µ tt; I~ a.. f.!: ·'¢( wm '(--, }« ~ 'r-r {;)· 0 0 >.::t ·,AWMWM ,A~ •M·,A •, 'OW %: w - Mi2J.1ma11:)d _.,,,~.., ~ 0 (.'>) N '"I ..,, 'T - - ,,., 0 -.:.:::-: ..... --· f~ ....... ~ ,~\;,j•t: r.?-!.if:~/? ~/:f:: ··, :.~.....'!'·";•,;-·.:·~:·:·~ .:-:;::~ lh.~-:;_/~ ~/~//~ 435 40;: 403 !: ~rsp.H~ t~xr:1r:=:-~~-ff>i:=:. t>:::-:-:.~u: ::-:• ~n>:::-~n:.ny I .j t;-;:u:}; ?:-;:GPU . Sl i~:~t1-:::-r~;: fi t)fn ~b:=:::!Jt:;::· r'f}f.::f::K::ry b::~nK:r:-GPU ;♦------------------ :r 480,----~--~ i Outp~:ttf:<_:,;·:sp<::, i vpk-.:;;:r!to ::';ix.~:Jrfs j r?>:3 r":::(·!~).. (3;~t: k ,...................... . 440 i::;:~~=~:,:~:;:::;t~:~::~ f 4~}fj V'.J::i~t fr_;r:~w::tp i fr:=!· \:\'°::~~t f)W{)p r,..· ✓ y5 O{ lnpcl!i◊tl!f,rn t:::•:-:.b.Jr<:: 44~=, ........ p:_);ntr~f::'> 1............. . t..l~~;~~;~r•:::;~t',J N.:.'::.................. L:~::-(t-..... .........., --------~~.......... tt-~~ratK>11?,,.""r·-.:-- ..... ........_ ......... ✓ Y<::-s Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 11 of 20 e • 00 • ~ ~ ~ ~ = ~ 3 .i 2 3 4 1 2 3 -··.. -·· ...··········· .............-...... · .... 1 2 3 ~--•-•.• -~···'····' 6 5 6 7 8 4 5 6 4 5 6 ~ ~ 7!8 9 ·g 7 8 9 7 8 9 :-: ' " .... ... ~ N 0 ::a) b) ........,.,., ,, ...... ' .... -·· )1 ······=·······--••'·•·· N ~ Bo1dlines shmv the packing of data into pixels ',Nlthfour color components, rJJ =- ('D .... ('D Ul .... 0 Ul FIG»5 d rJl. ~ ~ \0 ~ 0--, "'""' ~ Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 12 of 20 US RE49,461 E 1 2 GRAPHIC PROCESSOR BASED processing unit/system (CPU). Complete independence is ACCELERATOR SYSTEM AND METHOD not desirable, however; user input might affect how the computation is performed and even interrupt it if necessary. Furthermore, the user output and the disk output are depen- Matter enclosed in heavy brackets [ ] appears in the 5 dent on the results of the computation. A reasonable solution original patent but forms no part of this reissue specifica- would be to separate input/output into threads, so that it is tion; matter printed in italics indicates the additions interacting with hardware occurs in parallel with the com- made by reissue; a claim printed with strikethrough putation. In this case whatever CPU processing is required indicates that the claim was canceled, disclaimed, or held for input/output should be designed so that it provides the invalid by a prior post-patent action or proceeding. 10 synchronization with computation. In the case of GPGPU, the computation itself is performed RELATED APPLICATIONS outside of the CPU, so the complete system comprises three "peripheral" components: user interactive hardware, disk The present application is a reissue continuation appli- hardware, and computational hardware. The central process- cation of U.S. application Ser. No. 15/808,201, which was 15 ing unit (CPU) establishes communication and synchroni- filed on Nov. 9, 2017 as a broadening reissue application of zation between peripherals. Each of the peripherals is pref- U.S. Pat. No. 9,189,828, filed Jan. 3, 2014, which claims a erably controlled by a dedicated thread that is executed in priority benefit, under 35 U.S.C. §120, as a continuation of parallel with minimal interactions and dependencies on the U.S. application Ser. No. 11/860,254, now U.S. Pat. No. other threads. 8,648,867 B2, filed Sep. 24, 2007, entitled "Graphic Pro- 20 A GPU on a conventional video card is usually controlled cessor Based Accelerator System and Method," which in through OpenGL, DirectX, or similar graphic application turn claims the priority benefit, under 35 U.S.C. §119(e), of programming interfaces (APis ). Such APis establish the U.S. Application No. 60/826,892, filed Sep. 25, 2006. The context of graphic operations, within which all calls to the present application is also a broadening reissue application GPU are made. This context only works when initialized of U.S. Pat. No. 9,189,828, filed Jan. 3, 2014, which is a 25 within the same thread of execution that uses it. As a result, continuation of U.S. application Ser. No. 11/860,254, now in a preferred embodiment, the context is initialized within U.S. Pat. No. 8,648,867 B2, filed Sep. 24, 2007, which in a computational thread. This creates complications, how- turn claims the priority benefit, under 35 U.S. C. § 119(e), of ever, in the interaction between the user interface thread that U.S. Application No. 60/826,892, filed Sep. 25, 2006. Each changes parameters of simulations and the computational of the above-identified applications is incorporated herein by 30 thread that uses these parameters. reference in its entirety. More than one reissue application A solution as proposed here is an implementation of the has been filed for the reissue of U.S. Pat. No. 9,189,828, computational stream of execution in hardware, so that including this application and U.S. application Ser. No. thread and context initialization are replaced by hardware 15/808,201. initialization. This hardware implementation includes an 35 expansion card comprising a printed circuit board having (a) BACKGROUND one or more graphics processing units, (b) two or more associated memory banks that are logically or physically Graphics Processing Units (GPUs) are found in video partitioned, (c) a specialized controller, and (d) a local bus adapters (graphic cards) of most personal computers (PCs), providing signal coupling compatible with the PCI industry video game consoles, workstations, etc. and are considered 40 standards (this includes but is not limited to PCI-Express, highly parallel processors dedicated to fast computation of PCI-X, USB 2.0, or functionally similar technologies). The graphical content. With the advances of the computer and controller handles most of the primitive operations needed to console gaming industries, the need for efficient manipula- set up and control GPU computation. As a result, the CPU tion and display of 3D graphics has accelerated the devel- is freed from this function and is dedicated to other tasks. In opment of GPUs. 45 this case a few controls (simulation start and stop signals In addition, manufacturers of GPU shave included general from the CPU and the simulation completion signal back to purpose programmability into the GPU architecture leading CPU), GPU programs and input/output data are the infor- to the increased popularity of using GPU s for highly paral- mation exchanged between CPU and the expansion card. lelizable and computationally expensive algorithms outside Moreover, since on every time step of the simulation the of the computer graphics domain. When implemented on 50 results from the previous time step are used but not changed, conventional video card architectures, these general purpose the results are preferably transferred back to CPU in parallel GPU (GPGPU) applications are not able to achieve optimal with the computation. performance, however. There is overhead for graphics- In general, according to one aspect, the invention features related features and algorithms that are not necessary for a computer system. This system comprises a central pro- these non-video applications. 55 cessing unit, main memory accessed by the central process- ing unit, and a video system for driving a video monitor in SUMMARY response to the central processing unit as is common. The computer system further comprises an accelerator that uses Numerical simulations, e.g., finite element analysis, of input data from and provides output data to the central large systems of similar elements (e.g. neural networks, 60 processing unit. This accelerator comprises at least one genetic algorithms, particle systems, mechanical systems) graphics processing unit, accelerator memory for the graphic are one example of an application that can benefit from processing unit, and an accelerator controller that moves the GPGPU computation. During numerical simulations, disk input data into the at least one graphics processing unit and and user input/output can be performed independently of the accelerator memory to generate the output data. computation because these two processes require interac- 65 In the preferred, the central processing unit transfers the tions with peripheral hardware (disk, screen, keyboard, input data for a simulation to the accelerator, after which the mouse, etc) and put relatively low load on the central accelerator executes simulation computations to generate Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 13 of 20 US RE49,461 E 3 4 the output data, which is transferred to the central processing In more detail, the computer system 100 in one example unit. Preferably, the accelerator controller dictates an order is a standard personal computer (PC). However, this only of execution of instructions to the at least one graphics serves as an example environment as computing environ- processing unit. The use of the separate controller enables ment 100 does not necessarily depend on or require any data transfer during execution such that the accelerator 5 combination of the components that are illustrated and controller transfers output data from the accelerator memory described herein. In fact, there are many other suitable to main memory of the central processing unit. computing environments for this invention, including, but In the preferred embodiment, the accelerator controller not limited to, workstations, server computers, supercom- comprises an interface controller that enables the accelerator puters, notebook computers, hand-held electronic devices to communicate over a bus of the computer system with the 10 such as cell phones, mp3 players, or personal digital assis- central processing unit. tants (PDAs), multiprocessor systems, progranimable con- In general according to another aspect, the invention also sumer electronics, networks of any of the above-mentioned features an accelerator system for a computer system, which computing devices, and distributed computing environments comprises at least one graphics processing unit, accelerator that including any of the above-mentioned computing memory for the graphic processing unit and an accelerator 15 devices. controller for moving data between the at least one graphics In one implementation the GPU accelerator is imple- processing unit and the accelerator memory. In general according to another aspect, the invention also mented as an expansion card 180 includes connections with features a method for performing numerical simulations in a the motherboard 110, on which the one or more CPU's 120 computer system. This method comprises a central process- 20 are installed along with main, or system memory 130 and ing unit loading input data into an accelerator system from mass/non volatile data storage 140, such as hard drive or main memory of the central processing unit and an accel- redundant array of independent drives (RAID) array, for the erator controller transferring the input data to a graphics computer system 100. In the current example, the expansion processing unit with instructions to be performed on the card 180 communicates to the motherboard 110 via a local input data. The accelerator controller then transfers output 25 bus 190. This local bus 190 could be PCI, PCI Express, data generated by the graphic processing unit to the central PCI-X, or any other functionally similar technology (de- processing unit as output data. pending upon the availability on the motherboard 110). An The above and other features of the invention including external version GPU accelerator is also a possible imple- various novel details of construction and combinations of mentation. In this example, the external GPU accelerator is parts, and other advantages, will now be more particularly 30 connected to the motherboard 110 through USB-2.0, IEEE described with reference to the accompanying drawings and 1394 (Firewire), or similar external/peripheral device inter- pointed out in the claims. It will be understood that the face. particular method and device embodying the invention are The CPU 120 and the system memory 130 on the moth- shown by way of illustration and not as a limitation of the erboard 110 and the mass data storage system 140 are invention. The principles and features of this invention may 35 preferably independent of the expansion card 180 and only be employed in various and numerous embodiments without communicate with each other and the expansion card 180 departing from the scope of the invention. through the system bus 200 located in the motherboard 110. A system bus 200 in current generations of computers have BRIEF DESCRIPTION OF THE DRAWINGS bandwidths from 3.2 GB/s (Pentium 4 withAGTL+, Athlon 40 XP with EV6) to around 15 GB/s (Xeon Woodcrest with In the accompanying drawings, reference characters refer AGTL+, Athlon 64/Opteron with Hypertransport), while the to the same parts throughout the different views. The draw- local bus has maximal peak data transfer rates of 4 GB/s ings are not necessarily to scale; emphasis has instead been (PCI Express 16) or 2 GB/s (PCI-X 2.0). Thus the local bus placed upon illustrating the principles of the invention. Of 190 becomes a bottleneck in the information exchange the drawings: 45 between the system bus 200 and the expansion card 180. The FIG. 1 is a schematic diagram illustrating a computer design of the expansion card and methods proposed herein system including the GPU accelerator according to an minimizes the data transfer through the local bus 190 to embodiment of the present invention; reduce the effect of this bottleneck. FIG. 2 is block diagram illustrating the architecture for the The system memory 130 is referred to as the main GPU accelerator according to an embodiment of the present 50 random-access memory (RAM) in the description herein. invention; However, this is not intended to limit the system memory FIG. 3 is a block/flow diagram illustrating an exemplary 130 to only RAM technology. Other possible computer implementation of the top level control of the GPU accel- storage media include, but are not limited to ROM, erator system; EEPROM, flash memory, or any other memory technology. FIG. 4 is a flow diagram illustrating an exemplary imple- 55 In the illustrated example, the GPU accelerator system is mentation of the bottom level control of the GPU accelerator implemented on an expansion card 180 on which the one or system that is used to execute the target computation; and more GPU's 240 are mounted. It should be noted that the FIG. 5 is an example population of nine computational GPU accelerator system GPU 240 is separate from and elements arranged in a 3x3 square and a potential packing independent of any GPU on the standard video card 150 or scheme for texture pixels, according to an implementation of 60 other video driving hardware such as integrated graphics the present invention. systems. Thus the computations performed on the expansion card 180 do not interfere with graphics display (including DETAILED DESCRIPTION but not limited to manipulation and rendering of images). Various brand of GPU are relevant. Under current tech- FIG. 1 shows a computer system 100 that has been 65 nology, GPU's based on the GeForce series from NVIDIA constructed according to the principles of the present inven- Corporation or the Catalyst series from ATI/Advanced tion. Micro Devices, Inc. Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 14 of 20 US RE49,461 E 5 6 The output to a video monitor 170 is preferably through partition 250b is designed to hold the data textures repre- the video card 150 and not the GPU accelerator system 180. senting internal variables. The third partition 250c is The video card 150 is dedicated to the transfer of graphical designed to hold the data textures used as input at a information and connects to the motherboard 110 through a particular computation step on the GPU 240. The fourth local bus 160 that is sometimes physically separate from the 5 partition 250d holds the data textures used to accommodate local bus 190 that connects the expansion card 180 to the the output of a particular computational step on the GPU motherboard 110. 240. This partitioning scheme can be done logically, does FIG. 2 is a block diagram illustrating the general archi- not require hardware implementation. Also the partitioning tecture of the GPU accelerator system and specifically the scheme is also altered based on new designs or needs of the expansion card 180 in which at least one GPU 240 and 10 algorithms being employed. The reason for this partitioning associated memories 210 and 250 are mounted. Electrical is further explained in the Data Organization section, below. (signal) and mechanical coupling with a local bus 190 A local bus interface 230 on the controller 220 serves as provides signal coupling compatible with the PCI industry a driver that allows the controller 220 to communicate standards (this includes but is not limited to PCI, PCI-X, PCI through the local bus 190 with the system bus 200 and thus Express, or functionally similar technology). 15 the CPU 120 and RAM 130. This local bus interface 230 is The GPU accelerator further preferably comprises one not intended to be limited to PCI related technology. Other specifically designed accelerator controller 220. Depending drivers can be used to interface with comparable technology upon the implementation, the accelerator controller 220 is as a local bus 190. field programmable gate array (FPGA) logic, or custom built Data Organization application-specific (ASIC) chip mounted in the expansion 20 Each computational element discussed above has output card 180, and in mechanical and signal coupling with the variables that affect the rest of the system. For example in GPU 240 and the associated memories 210 and 250. During the case of a neural network it is the output of a neuron. A initial design, a controller can be partially or even fully computational element also usually has several internal implemented in software, in one example. variables that are used to compute output variables, but are The controller 220 commands the storage and retrieval of 25 not exposed to the rest of the system, not even to other arrays of data (on a conventional video card the arrays of elements of the same population, typically. Each of these data are represented as textures, hence the term 'texture' in variables is represented as a texture. The important differ- this document refers to a data array unless specified other- ence between output variables and internal variables is their wise and each element of the texture is a pixel of color access. information), execution of GPU programs (on a conven- 30 Output variables are usually accessed by any element in tional video card these programs are called shaders, hence the system during every time step. The value of the output the term 'shader' in this document refers to a GPU program variable that is accessed by other elements of the system unless specified otherwise), and data transfer between the corresponds to the value computed on the previous, not the system bus 200 and the expansion card 180 through the local current, time step. This is realized by dedicating two textures bus 190 which allows communication between the main 35 to output variables----one holds the value computed during CPU 120, RAM 130, and disk 140. the previous time step and is accessible to all computational Two memory banks 210 and 250 are mounted on the elements during the current time step, another is not acces- expansion card 180. In some example, these memory banks sible to other elements and is used to accumulate new values separated in the hardware, as shown, or alternatively imple- for the variable computed during the current time step. mented as a single, logically partitioned memory compo- 40 In-between time steps these two textures are switched, so nent. that newly accumulated values serve as accessible input The reason to separate the memory into two partitions 210 during the next time step, while the old input is replaced with 250 stems from the nature of the computations to which the new values of the variable. This switch is implemented by GPU accelerator system is applied. The elements of com- swapping the address pointers to respective textures as putation (computational elements) are characterized by a 45 described in the System and Framework section. single output variable. Such computational elements often Internal variables are computed and used within the same include one or more equations. Computational elements are computational element. There is no chance of a race con- same or similar within a large population and are computed dition in which the value is used before it is computed or in parallel. An example of such a population is a layer of after it has already changed on the next time step because neurons in an artificial neural network (ANN), where all 50 within an element the processing is sequential. Therefore, it neurons are described by the same equation. As a result, is possible to render the new value of internal variable into some data and most of the algorithms are common to all the same texture where the old was read from in the texture computational elements within population, while most of the memory bank. Rendering to more than one texture from a data and some algorithms are specific for each equation. single shader is not implemented in current GPU architec- Thus, one memory, the shader memory bank 210, is used to 55 tures, so computational elements that track internal variables store the shaders needed for the execution of the required would have to have one shader per variable. These shaders computations and the parameters that are common for all can be executed in order with internal variables computed computational elements and is coupled with the controller first, followed by output variables. 220 only. The second memory, the texture memory bank Further savings of texture memory is achieved through 250, is used to store all the necessary data that are specific 60 using multiple color components per pixel (texture element) for every computational element (including, but not limited to hold data. Textures can have up to four color components to, input data, output data, intermediate results, and param- that are all processed in parallel on a GPU. Thus, to eters) and is coupled with both the controller 220 and the maximize the use of GPU architecture it is desirable to pack GPU 240. the data in such a way that all four components are used by The texture memory bank 250 is preferably further par- 65 the algorithm. Even though each computational element can titioned into four sections. The first partition 250 a is have multiple variables, designating one texture pixel per designed to hold the external input data patterns. The second element is ineffective because internal variables require one Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 15 of 20 US RE49,461 E 7 8 texture and output variables require two textures. Further- Stream 303-runs on the GPU accelerator of the expansion more, different element types have different numbers of card 180 and interacts with the User Interaction Stream 302 variables and unless this number is precisely a multiple of through initialization routines and data exchange in between four, texture memory can be wasted. simulations. The Computational Stream 303 interacts with Amore reasonable packing scheme would be to pack four 5 the User Interaction Stream and the Data Output Stream computational elements into a pixel and have separate through synchronization procedures during simulations. textures for every variable associated with each computa- The crucial feature of the interaction between the User tional element. In this case the packing scheme is identical Interaction Stream 302 and the Computational Stream 303 is for all textures, and therefore can be accessed using the same the shift of priorities. Outside of the simulation, the system algorithm. Several ways to approach this packing scheme 10 100 is driven by the user input, thus the User Interaction are outlined here. An example population of nine computa- Stream 302 has the priority and controls the data exchange tional elements arranged in a 3x3 square (FIG. Sa) can be 304 between streams. After the user starts the simulation, the packed by element (FIG. Sb), by row (FIG. Sc), or by square Computational Stream 303 takes the priority and controls (FIG. Sd). the data exchange between streams until the simulation is Packing by element (FIG. Sb) means that elements 1,2,3,4 15 finished or interrupted 350. go into first pixel; 5,6,7,8 go into second pixel; 9 goes into The user starts 300 the framework through the means of third pixel. This is the most compact scheme, but not an operating system and interacts with the software through convenient because the geometrical relationship is not pre- the user interaction section 305 of the graphic user interface served during packing and its extraction depends on the size 306 executed on the CPU 120. The start 300 of the imple- of the population. 20 mentation begins with a user action that causes a GUI Packing by row (colunm; FIG. Sc) means that elements initialization 307, Disk input/output initialization 308 on the 1,2,3 go into pixel (1,1); 3,4,5 go into pixel (2,1), 7,8,9 go CPU 120, and controller initialization 320 of the GPU into pixel (3,1). With this scheme the element's y coordinate accelerator on the expansion card 180. GUI initialization in the population is the pixel's y coordinate, while the includes opening of the main application window and setting element's x coordinate in the population is the pixel's x 25 the interface tools that allow the user to control the frame- coordinate times four plus the index of color component. work. Disk I/O initialization can be performed at the start of Five by five populations in this case will use 2x5 texture, or the framework, or at the start of each individual simulation. 10 pixels. Five of these pixels will only use one out of four The user interaction 305 controls the setting and editing of components, so it wastes 37.5% of this texture. 25xl popu- the computational elements, parameters, and sources of lation will use 6xl texture (six pixels) and will waste 12.5% 30 external inputs. It specifies which equations should have of it. their output saved to disk and/or displayed on the screen. It Packing by square (FIG. Sd) means that elements 1,2,4,5 allows the user to start and stop the simulation. And it go into pixel (1,1); 3,6 go into pixel (1,2); 7,8 go into pixel performs standard interface functions such as file loading (2,1), and 9 goes into pixel (2,2). Both the row and the and saving, interactive help, general preferences and others. colunm of the element are determined from the row (col- 35 The user interaction 305 directs the CPU 120 to acquire unm) of the pixel times two plus the second (first) bit of the the new external input textures needed (this includes but is color component index. Five by five populations in this case not limited to loading from disk 140 or receiving them in will use 3x3 texture, or 9 pixels. Four of these pixels will real time from a recording device), parses them if necessary only use two out of four components, and one will only use 309, and initializes their transfer to the expansion card 180, one component, so it wastes 34.4% of this texture. This is 40 where they are stored 325 in the texture memory bank 250 more advantageous than packing by row, since the texture is by the controller 220. The user interaction 305 also directs smaller and the waste is also lower. 25xl population on the the CPU 120 to parse populations of elements that will be other hand will use 13xl texture (thirteen pixels) and waste used in the simulation, convert them to GPU programs >50% of it, which is much worse than packing by row. (shade rs), compile them 310, and initializes their transfer to In order to eliminate waste altogether the population 45 the expansion card 180, where they are stored 326 in the should have even dimensions in the square packing, and it shader memory bank 210 by the controller 220. This opera- should have a number of columns divisible by four in row tion is accompanied by the upload 309 of the initial data into packing. Theoretically, the chances are approximately the input partition of the texture memory bank 250, and equivalent for both of these cases to occur, so the particular stores the shader order of execution in the controller 220. task and data sizes should determine which packing scheme 50 The user can perform operations 309 and 310 as many times is preferable in each individual case. as necessary prior to starting the simulation or between The System and Framework simulations. FIG. 3 shows an exemplary implementation of the top The editing of the system between simulations is difficult level system and method that is used to control the compu- to accomplish without the hardware implementation of the tation. It is a representation of one of several ways in which 55 computational thread suggested herein. The system of equa- a system and method for processing numerical techniques tions (computational elements) is represented by textures can be implemented in the invention described herein and so that track variables plus shaders that define processing the implementation is not intended to be limited to the algorithms. As mentioned above, textures, shaders and other following description and accompanying figure. graphics related constructs can only be initialized within the The method presented herein includes two execution 60 rendering context, which is thread specific. Therefore tex- streams that run on the CPU 120-User Interaction Stream tures and shaders can only be initialized in the computa- 302 and Data Output Stream 301. These two streams pref- tional thread. erably do not interact directly, but depend on the same data Network editing is a user-interactive process, which accumulated during simulations. They can be implemented according to the scheme suggested above happens in the as separate threads with shared memory access and executed 65 User Interaction Stream 302. The simulation software thus on different CPUs in the case of multi-CPU computing has to take the new parameters from the User Interaction environment. The third execution stream-Computational Stream 302, communicate them to the Computational Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 16 of 20 US RE49,461 E 9 10 Stream 303 and regenerate the necessary shaders and tex- practical software implementation of the method and archi- tures. This is hard to accomplish without a hardware imple- tecture described above and pictorially represented in FIG. mentation of the Computational Stream 303. The Compu- 3. tational Stream 303 is forked from the User Interaction To use SANNDRA, the application should create a Stream and it can access the memory of the parent thread, 5 TSimulator object either directly or through inheritance. but the reverse communication is harder to achieve. The This object will handle global simulation properties and controller 220 allows operations 309 and 310 to be per- control the User Interaction Stream, Data Output Stream, formed as many times as necessary by providing the nec- and Computational Stream. Through TSimulator: :time- essary communication to the User Interaction Stream 302. step( ) TSimulator: :outfileinterval( ), and TSimulator: :out- After execution of the input parser texture generation 309 lO mode( ), the application can set the time step of the simu- and population parser shader generator and compiler 310 are lation, the time step of disk output, and the mode of the disk performed at least once, the user has the option to initialize output. The external input pattern should be packed into a the simulation 311. During this initialization the main con- TPattem object and bound to the simulation object through trol of the framework is transferred to the GPU accelerator 15 TSimulator::resetinputs( ) method. TSimulator::sim- system's accelerator controller 220 and computation 330 is Length( ) sets the length of the simulation. started (see FIG. 4; 420). The user retains the ability to The second step is to create at least one population of interrupt the simulation, change the input, or to change the equations (Tpopulation object). Population holds one equa- tion object TEquation. This object contains only a formula display properties of the framework, but these interactions and does not hold element-specific data, so all elements of are queued to be performed at times determined by the 20 the population can share single TEquation. controller-driven data exchange 314 and 316 to avoid the The TEquation object is converted to a GPU program corruption of the data. before execution. GPU programs have to be executed within The progress monitor 312 is not necessary for perfor- a graphical context, which is stream specific. TSimulator mance, but adds convenience. It displays the percentage of creates this context within a Computational Stream, there- completed time steps of the simulation and allows the user 25 fore all programs and data arrays that are necessary for to plan the schedule using the estimates of the simulation computation have to be initialized within Computational wall clock times. Controller-driven data exchange 314 Stream. Constructor of TPopulation is called from User updates the display of the results 313. Online screen output Interaction Stream, so no GPU-related objects can be ini- for the user selected population allows the user to monitor tialized in this constructor. the activity and evaluate the qualitative behavior of the 30 TPopulation: :fillElements( ) is a virtual method designed network. Simulations with unsatisfactory behavior can be to overcome this difficulty. It is called from within the terminated early to change parameters and restart. Control- Computational Stream after TSimulator: :networkCreate( ) is ler-driven data exchange 314 also drives the output of the called in the User Interaction Stream. A user has to override results to disk 317. Data output to disk for convenience can 35 TPopulation::fillElements( ) to create TEquation and other be done on an element per file basis. A suggested file format computation related objects both element independent and includes a leftmost colunm that displays a simulated time for element-specific. Element independent objects include sub- each of the simulation steps and subsequent colunms that components of TEquation and objects that describe how to display variable values during this time step in all elements handle interdependencies between variables implemented through derivatives of TGate class. with identical equations (e.g. all neurons in a layer of a 40 Element-specific data is held in TElement objects. These neural network). objects hold references to TEquation and a set of TGate Controller-driven data exchange or input parser texture objects. There is one TElement per population, but the size generator 316 allows the user to change input that is gen- of data arrays within this object corresponds to population erated on the fly during the simulation. This allows the size. All TElement objects have to be added to the TSimu- framework monitoring of the input that is coming from a 45 lator list of elements by calling TSimulator::addUnit( ) recording device (video camera, microphone, cell recording method from TPopulation: :fillElements( ). electrode, etc) in real time. Similar to the initial input parser Finally, TPopulation::fillElements() should contain a set 309, it preprocesses the input into a universal format of the of TElement::add*Dependency( ) calls for each element. data array suitable for texture generation and generates Each of these calls sets a corresponding dependency for textures. Unlike the initial parser 309, here the textures are 50 every TGate object. Here TGate object holds element inde- transferred to hardware not whenever ready but upon the pendent part of dependency and TElement:: request of the controller 220. add*Dependency( ) sets element-specific details. The controller 220 also drives the conditional testing 315 System provided TPopulation handles the output of com- and 318 informs the CPU-bound streams whether the simu- putational elements, both when they need to exchange the lation is finished. If so, the control returns to the User 55 data and when they need to output it to disk. User imple- Interaction Stream. The user then can change parameters or mentation of TPopulation derivative can add screen output. inputs (309 and 310), restart the simulation (311) or quit the Listing 1 is an example code of the user program that uses framework (390). a recurrent competitive field (RCF) equation: SANNDRA (Synchronous Artificial Neuronal Network Distributed Runtime Algorithm; http://www.kinness.net/ 60 LISTING 1 Docs/SANNDRA/html) was developed to accelerate and optimize processing of numerical integration of large non- uint16_t w - 3, h - 3; homogenous systems of differential equations. This library static float m_compet = 0.5; static float m_persist = 1.0; is fully reworked in its version 2.x.x to support multiple class TCablePopRCF : public TPopulation computational backends including those based on multicore 65 { CPUs, GPUs and other processing systems. GPU based TEq_RCF* m_equation; backend for SANNDRA-2.x.x can serve as an example Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 17 of 20 US RE49,461 E 11 12 LISTING I-continued for new time step. To avoid data confusion, the new values of variables should be rendered in a separate texture. After TGate* m_gatel; the time step is completed for all equations, these new values TGate* m_gate2; void createGatingStructure( ) should be copied over old values so that they are used as { 5 input during the next time step. Copying textures is an m_gatel - new TGate(0); expensive operation, computationally, but since the textures m_gate2 - new TGate(l); are referred to by texture IDs (pointers), swapping these }; pointers for input and output textures after each time step void createUnitStructure(TBasicUnit* u) achieves the same result at a much lesser cost. { u->addO2OPlnputDependency(m_gatel, 0., 0., 0.004, 0., 0, 0); 10 In the hardware solution suggested herein, ID swapping is u->addFullDependency(m_gate2, population()); equivalent to swapping the base memory address for two } partitions of the texture memory bank 250. They are public: TCablePopRCF() : TPopulation("compCPU RCF", w, h, true) { }; swapped 485 during synchronization (485, 430, and 455) so ~TCablePopRCF() {if(m_equation) delete m_equation; that data transfer 445 and the computation 435-487 proceeds if(m_gatel) delete m_gatel; if(m_gate2) delete m_gate2;}; immediately and in parallel with data transfer as shown in 15 FIG. 4. A hardware solution allows this parallelism through boo! fillElements(TSimulator* sim); }; access of the controller 220 to the onboard texture memory boo! TCablePopRCF::fillElements(TSimulatior* sim) bank 250. { The main computation and data exchange are executed by m_equation - new TEq_RCF(this, m_compet, m_persist); createGatingStructure( ); the controller 220. It runs three parallel substreams of for(size_t i - 0; i < xSize( ); ++i) 20 execution: Computational Substream 403, Data Output Sub- for(size_t j - 0; j < ySize( ); ++j) stream 402, and Data Input Substream 404. These streams { are synchronized with each other during the swap of pointers TElement* u - new TCPUElement(this, m_equation, i, j); 485 to the input and output texture memory partitions of the sim->addUnit(u); create U nitStructure( u); texture memory bank 250 and the check for the last iteration } 25 487. Algorithmically, these two operations are a single Return true; atomic operation, but the block diagram shows them as two } separate blocks for clarity. int The Computational Substream 403 performs a computa- main() tional cycle including a sequential execution of all shaders { // Input pattern generation (309 in FIG.3) that were stored in the shader memory bank 210 using the 30 uint32_t* pat - new uint32_t[w*h]; appropriate input and output textures. To begin the simula- TRandom randGen (0); tion the controller 220 initializes three execution sub streams for(uint32_t I - 0; I < w*h; ++i) 403, 402, and 404. On every simulation step, the Compu- pat[i] - randGen.random( ); tational Substream 403 determines which textures the GPU Tpattern* p - new Tpattern(pat, w, h); // Setting up the simulation 240 will need to perform the computations and initiates the 35 upload 435 of them onto the GPU 240. The GPU 240 can TSimulator* cableSim - new TSimulator("data"); //(308 and 320 in FIG. 3) communicate directly with the texture memory bank 250 to cableSim->timestep(0.05); //(320 in FIG. 3) upload the appropriate texture to perform the computations. cableSim->resetlnputs(p); //(325 in FIG. 3) The controller 220 also pulls the first shader (known by the cableSim->outfileinterval(0.1); //(308 in FIG. 3) cableSim->outmode(SANNDRA::timefunc); //(308 in FIG. 3) stored order) from the shader memory bank 210 and uploads cableSim->simLength(60.0); //(320 in FIG. 3) 40 450 it onto the GPU 240. // Preparing the population The GPU 240 executes the following operations in this TPopulation* cablePop - new TCablePopRCF( ); //(310 in FIG. 3) order: performs the computation (execution of the shader) cableSim->networkCreate( ); //(326 in FIG. 3) 470; tells the controller 220 that it is done with the compu- uintl 6_t user= 1; tations for the current shader; and after all shaders for this while(user) { 45 particular equation are executed sends 480 the output tex- if(! cableSim->simulationStart(true, 1)) //(311 in FIG. 3) tures to the output portion of the texture memory bank 250. exit(!); This cycle continues through all of the equations based on std::cout<<"Repeat?ln"; //(305 in FIG. 3) the branching step 482. std::cin>>user; //(305 in FIG. 3) An example shader that performs fourth order Runge- if(user -- 1) cableSim->networkReset( ); //(305 in FIG. 3) Kutta numerical integration is shown in Listing 2 using 50 { GLSL notation; If(cableSim) Delete cableSim; //Also deletes cablePop and its internals LISTING 2 exit(0); }; uniform sarnpler2DRect Variable; 55 uniform float integration_step; float halfstep - integration_step*0.5; FIG. 4 is a detailed flow diagram illustrating a part of an float fl_6step - integration_step/6.0; exemplary implementation of the bottom level system and vec4 output - texture2DRect(Variable, gl_TexCoord[0].st); // define equation( ) here method performed during the computation on the GPU vec4 rungekutta4(vec4 x) accelerator of the expansion card 180 and is a more detailed { view of the computational box 330 in FIG. 3. FIG. 4 is a 60 canst vec4 kl - equation(x); representation of one of several ways in which a system and canst vec4 k2 - equation(x + halfstep*kl); method for processing numerical techniques can be imple- canst vec4 k3 - equation(x + halfstep*k2); canst vec4 k4 - equation(x + integration step*k3); mented. return fl_6step*(kl + 2.0*(k2 + k3) + k4); With systems of equations that have complex interdepen- } dencies it is likely that the variable in some equation from 65 Void main(void) a previous time step has to be used by some other equation { after the new values of this variable are already computed Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 18 of 20 US RE49,461 E 13 14 LISTING 2-continued cycle (or less frequently as defined by the user), writing this output to disk 140, and displaying this output on the monitor output +- rungekutta4( output); 170. This frees the CPU 120 to execute other applications gl_FragColor - output; } and allows the expansion card to run at its full capacity 5 without being slowed down by extensive interactions with the CPU 120. The shader in Listing 2 can be executed on conventional 2. Minimizing data transfer between the expansion card video card. Using the controller 220 this code can be further 180 and the system bus 200. All of the information needed optimized, however. Since the integration step does not to perform the simulations will be stored on the expansion change during the simulation, the step itself as well as the 10 card 180 and all simulations will take place on it. Further- halfstep and 1/4 of the step can be computed once per more, whatever data transfer remains necessary will take simulation, and updated in all shaders by a shader update place in parallel with the computation, thus reducing the procedures 310, 326 discussed above. impact of this transfer on the performance. After all of the equations in the computational cycle are 3. New way to execute GPU programs (shaders). Previ- computed the main execution substream 403 on the control- 15 ously, the CPU 120 had full control over the order of !er 220 can switch 485 the reference pointers of the input and shader's execution and was required to produce specific output portions of the texture memory bank 250. commands on every cycle to tell the GPU 240 which shader The two other substreams of execution on the controller to use. With the invention disclosed herein, shaders will 220 are waiting (blocks 430 and 455, respectively) for this initially be stored on the shader memory bank 210 on the switch to begin their execution. The Data Input Substream 20 expansion card 180 and will be sent to the GPU 240 for 404 is controlling 440 the input of additional data from the execution by the general purpose controller 220 located on CPU 120. This is necessary in cases where the simulation is the expansion card. monitoring the changing input, for example input from a 4. Multiple parallelisms. The GPU 240 is inherently video camera or other recording device in the real time. This parallel and is well suited to perform parallel computations. substream uploads new external input from the CPU 120 to In parallel with the GPU 240 performing the next calcula- the texture memory bank 250 so it can be used by the main 25 computational sub stream 403 on the next computational step tion, the controller 220 is uploading the data from the and waits for the next iteration 475. The Data Output previous calculation into main memory 130. Furthermore, Substream 445 controls the output of simulation results to the CPU 120 at the same time uses uploaded previous results the CPU 120 if requested by the user. This substream to save them onto disk 140 and to display them on the screen uploads the results of the previous step to the main RAM 30 through the system bus 200. 130 so that the CPU 120 can save them on disk 140 or show 5. Reuse of existing and affordable technology. All hard- them on the results display 313 and waits for the next ware used in the invention and mentioned here-in are based iteration 460. on currently available and reliable components. Further Since the Computational Substream 403 determines the advance of these components will provide straightforward timing of input 440 and output 445 data transfers, these data 35 improvements of the invention. transfers are driven by the controller 220. To further reduce While this invention has been particularly shown and the data transfer overhead (and disk 140 overhead also) the described with references to preferred embodiments thereof, controller 220 initiates transfer only after selected compu- it will be understood by those skilled in the art that various tational steps. For example, if the experimental data that is changes in form and details may be made therein without simulated was recorded every 10 milliseconds (msec) and 40 departing from the scope of the invention encompassed by the simulation for better precision was computed every 1 the appended claims. msec, then only every tenth result has to be transferred to match the experimental frequency. This solution stores two copies of output data, one in the What is claimed is: expansion card texture memory bank 250 and another in the [1. A computer system, comprising: system RAM 130. The copy in the system RAM 130 is 45 a central processing unit to receive input data; accessed twice: for disk I/O and screen visualization 313. An main memory, operably coupled to the central processing alternative solution would be to provide CPU 120 with a unit via a bus, to store the input data received by the direct read access to the onboard texture memory bank 250 central processing unit; by mapping the memory of the hardware onto a global an accelerator, operably coupled to the central processing memory space. The alternative solution will double the 50 unit and the first memory via the bus, to receive at least communication through the local bus 190. Since the goal a portion of the input data from the main memory, the discussed herein is reducing the information transfer through accelerator comprising: the local bus 190, the former solution is favored. at least one graphics processing unit to perform a The main substream 403 determines if this is the last sequence of computations on the at least a portion of iteration 487. If it is the last iteration, the controller 220 55 the input data so as to generate output data, inter- waits for the all of the execution substreams to finish 490 mediate computations in the sequence of computa- and then returns the control to the CPU 120, otherwise it tions yielding intermediate results; and begins the next computational cycle. accelerator memory, operably coupled to the graphic This repeats through all of the computational cycles of the processing unit, to store the results of the plurality of simulation. 60 sequential computations; and CONCLUSION a controller, operably coupled to the at least one graphics processing unit and the accelerator memory, to transfer This GPU accelerator system offers the following poten- the at least a portion of the input data into the accel- tial advantages: erator memory, and to transfer at least a portion of the 1. Limited computations on the CPU 120. The CPU 120 65 output data from the accelerator memory to the main is only used for user input, sending information to the memory during performance of the sequence of com- controller 220, receiving output after each computational putations by the at least one graphic processing unit.] Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 19 of 20 US RE49,461 E 15 16 [2. The computer system of claim 1, wherein the central [13. The method of claim 12, further comprising: processing unit is configured to receive the input data in storing the input data in the main memory in response to response to a user interaction.] a user interaction.] [3. The computer system of claim 1, wherein: [14. The method of claim 12, further comprising: the central processing unit is configured to receive the 5 receiving the input data at a first rate; and input data at a first rate; and wherein (A) comprises performing the sequence of com- the at least one graphics processing unit is configured to putations at a second rate different than the first rate.] perform the sequence of computations at a second rate [15. The method of claim 12, wherein (A) comprises: different than the first rate.] generating an output representative of an output of at least 10 [4. The computer system of claim 1, wherein the main one neuron in an artificial neural network.] memory is configured to store a copy of the output data [16. The method of claim 12, wherein (C) comprises: stored in the accelerator memory.] transferring the second portion of the output data from the [5. The computer system of claim 1, wherein an output of accelerator memory to the main memory without trans- at least one computation in the sequence of computations 15 ferring any of the intermediate results of the plurality of represents an output of at least one neuron in an artificial sequential computations from the accelerator memory neural network.] to the main memory so as to reduce data transfer via the [6. The computer system of claim 1, wherein accelerator bus.] memory comprises: [17. The method of claim 12, wherein (C) comprises: a first memory bank to store parameters common to all of 20 transferring the second portion of the output data from the the computations in the sequence of computations; and accelerator memory to the main memory after the GPU a second memory bank to store data specific to at least one has begun to perform another sequence of computa- computation in the sequence of computations.] tions.] [7. The computer system of claim 1, wherein the control- [18. The method of claim 17, wherein (C) further com- ler is configured to transfer the output data from the accel- 25 prises: erator memory to the main memory without transferring any initiating transfer of the second portion of the output data of the intermediate results from the accelerator memory to in parallel with performance of at least one computa- the main memory so as to reduce data transfer via the bus.] tion in the other sequence of computations.] [8. The computer system of claim 1, wherein the control- [19. The method of claim 12, further comprising: ler is configured to transfer at least a portion of the output 30 acquiring the input data in real time with at least one of data from the accelerator memory to the main memory after a video camera, a microphone, or a cell recording the at least one graphics processing unit has begun to electrode operably coupled to the CPU.] perform another sequence of computations.] [20. The method of claim 12, further comprising: [9. The computer system of claim 8, wherein the control- 35 storing parameters common to all of the computations in !er is configured to initiate transfer of the at least a portion the sequence of computations in a first memory bank in of the input data and to transfer the at least a portion of the the accelerator memory; and output data in parallel with performance of at least one storing data specific to at least one computation in the computation in the other sequence of computations by the at sequence of computations in a second memory bank in least one graphics processing unit.] 40 the accelerator memory.] [10. The computer system of claim 1, wherein the con- 21. A method of executing computations representing an troller is configured to control execution of the sequence of artificial neural network on a computer system comprising computations by the at least one graphics processing unit.] at least one central processing unit (CPU), a processing [11. The computer system of claim 1, further comprising: unit, a first memory partition, and a second memory parti- at least one of a video camera, a microphone, or a cell 45 tion, the method comprising: recording electrode, operably coupled to the central executing, by the at least one CPU, a user interaction processor unit, to acquire the input data in real time.] stream, the user interaction stream controlling transfer [12. A method of performing a sequence of computations of inputs to the artificial neural network to the first on a computer system comprising a central processing unit memory partition and the second memory partition; (CPU), a main memory operably coupled to the central 50 executing, by the processing unit, a computational stream, processing unit via a bus, an accelerator operably coupled to the computational stream controlling data exchange the CPU and the main memory via the bus, the accelerator between the user interaction stream and the computa- comprising a graphics processing unit (GPU) and an accel- tional stream during execution of the computations erator memory, the method comprising: representing the artificial neural network; (A) performing, by the GPU, the sequence of computa- 55 shifting control of a data exchange between the user tions on a first portion of the input data so as to generate interaction stream and the computational stream to the a first portion of the output data, intermediate compu- computational stream in response to starting execution tations in the sequence of computations yielding inter- of the computations representing the artificial neural mediate results; network; (B) in parallel with performing the sequence of compu- 60 shifting control of the data exchange between the user tations by the GPU in (A), transferring a second portion interaction stream and the computational stream to the of the input data from the main memory to the accel- user interaction stream in response to completion or erator via the bus; and interruption of the computations representing the arti- (C) in parallel with performing the sequence of compu- ficial neural network; tations by the GPU in (A), transferring a second portion 65 queueing a user command received by the user interaction of the output data from the accelerator memory to the stream during execution of the computations represent- main memory via the bus.] ing the artificial neural network; and Case 7:26-mc-00318-LS Document 6-5 Filed 08/18/26 Page 20 of 20 US RE49,461 E 17 18 executing the user command during execution of the a second memory partition; computations representing the artificial neural network at least one central processing unit (CPU), operably at times determined by the computational stream. coupled to the camera, the first memory partition, and 22. The method of claim 21, wherein the user interaction the second memory partition, to execute a user inter- stream controls the data exchange between the user inter- 5 action stream, the user interaction stream controlling action stream and the computational stream outside of transfer of the input data acquired by the camera to the execution of the computations representing the artificial first memory partition and the second memory partition neural network. during execution of the computations representing the 23. The method of claim 21, wherein executing the user artificial neural network; interaction stream comprises: 10 a processing unit, operably coupled to the first memory controlling setting and editing of computational elements partition, the second memory partition, and the at least of the computations representing the artificial neural one CPU, to execute a computational stream, the network 24. The method of claim 21, wherein executing the user computational stream controlling transfer of the input interaction stream comprises: 15 data from the first memory partition and the second controlling setting and editing of parameters of the com- memory partition during execution of the computations putations representing the artificial neural network. representing the artificial neural network, the execution 25. The method of claim 21, wherein executing the user of the computations representing the artificial neural interaction stream comprises: network occurring while the camera is acquiring the controlling setting and editing of parameters of the inputs 20 input data; and to the artificial neural network. a controller, operably coupled to the at least one CPU and 26. The method of claim 21, wherein executing the user the processing unit, to queue user interactions received interaction stream comprises: by the user interaction stream during the execution of specifying an output to be saved to disk and/or displayed the computations representing the artificial neural net- on a screen. 25 work for performance at times selected to avoid data 27. The method of claim 21, wherein executing the user corruption. interaction stream comprises: 34. The system of claim 33, wherein the user interactions parsing elements to be used in the computations repre- cause interruption of the computations representing the senting the artificial neural network. artificial neural network. 28. The method of claim 27, wherein the processing unit 30 35. The system of claim 33, wherein the user interactions comprises a graphics processing unit ( GPU) and executing cause a change in inputs to the artificial neural network. the user interaction stream further comprises: 36. The system of claim 33, wherein the user interactions converting the elements into GPU programs. cause a change in display properties of an output of the 29. The method of claim 28, wherein executing the user computations representing the artificial neural network. interaction stream comprises: 35 3 7. The system of claim 33, wherein the controller is compiling the GPU programs. configured to request the input data during the execution of 30. The method of claim 29, wherein executing the user the computations representing the artificial neural network. interaction stream comprises: 38. The method of claim 21, wherein the user command transferring the GPU programs to the second memory causes interruption of the computations representing the partition. 40 artificial neural network. 31. The method of claim 21, further comprising: 39. The method of claim 21, wherein the user command executing, by the at least one CPU, a data output stream, causes a change in the inputs to the artificial neural net- the data output stream controlling transfer of outputs of work. the computations representing the artificial neural net- 40. The method of claim 21, wherein the user command work to disk. 45 causes a change in display properties of an output of the 32. The method of claim 21, further comprising: computations representing the artificial neural network. generating the inputs with a video camera during execu- 41. The method of claim 21, wherein, during execution of tion of the computations. the computations representing the artificial neural network, 33. A system for executing computations representing an the computational stream controls the data exchange artificial neural network, the system comprising: 50 between the user interaction stream and the computational a camera to acquire input data for the artificial neural stream by requesting the inputs to the artificial neural network; network. a first memory partition; * * * * *