My generative programming Toolkit has two separate interpreters in it, one that is used to parse the input code into a data structure, and one which is used to merge information from that data structure with some template code to create the output. Both languages are similar in the minor details, but the special features which make them work well for their specific purpose are quite different. Earlier versions of these languages used completely separate interpreter code, which had a very close alignment with the syntax of the language. This was not really a very good approach, as it meant that minor changes in the language syntax could require significant changes in the interpreter, and what shared code there was was largely done by cutting and pasting rather than a common code base. It also meant that any future attempt to translate the PHP prototype code into a more conventional language for production use would involve a lot of duplication of effort.
The new versions of the languages take a lot of ideas from the previous versions, but the are complete rewrites. Rather than writing two new interpreters, or trying to modify the previous interpreters, which would probably have been a bigger job than writing new ones, I decided to change the architecture slightly. The new plan involves a slightly lower level interpreter, with some aspects of a virtual machine, which is designed as a core that supports the basic common structures of conventional programming languages, that can also be extended to support the special features of the individual languages.
To start this project I created a very simple “core language”, which has the basic structures such as variables, arrays, functions and blocks used in any conventional programming language. It doesn’t have features that my interpreters do not need such as floating-point numbers, and it doesn’t have any features that are remotely unique. All that was needed at this point was a language which was sufficient to allow me to create some simple examples for development purposes.
I used my existing Helper tool to create a parser for this ‘core’ language, and then generated a PHP implementation using its export function. PHP parsers created by that tool need a ‘builder’ class created to actually process the parse matches into a useful data-structure, in much the same way as ANTLR 4. Helper creates a skeleton, but the details need filled in by hand. This is the point where the actual design of my reusable interpreter core started.
I decided some time ago that all data passed between different parts of my toolchain should be in the form of JSON files, because using a standard well structured format will make it easier to port the tools to a more appropriate language for production use in stages at a later date. (PHP is an excellent language for prototyping this type of tool, as it has got good support for text processing, and faster regular expression matching than most compiled languages, however I doubt many end-users would wish to set up a PHP environment just to run these tools.)
The JSON code above shows the outer for loop from the first example in its semi-compiled form. The variables init, test and step contain the pseudo-assembler for instances of the tiny stack-based virtual machine. The content of the instructions variable has been removed, but this would contain all the instructions contained within the for loop’s block. My JSON representation of the input retains all the original structure in a fairly clear form, but translates it into a much simpler representation for an interpreter. In some ways this is a similar approach to the response processing section of the QTI 2 item, though when we were developing that we had to be very careful not to use any terminology similar to compilers or virtual machines for political reasons.
In the Java virtual machine, stack frames are created whenever a new method is entered. To some extent this is similar to my use of new tiny virtual machines for each expression, but as the JVM operates at a much lower level, these frames are created as a higher level than in my system. That helps provide support for scope. In my CoreLang scope handling is provided through the data access class which handles storage of variables. Scope is controlled by instructions in the JSON language, which means that the CoreLang can cope with input with and without block scope (or even with no meaningful scope like early versions of BASIC.) My parser and template languages have slightly different requirements – in the parser language handling of variables and scope is conventional, however in the template language, data that is passed in from the parser is read only. This means the data classes will be slightly different.
The Core Language is implemented as an extensible class, and makes use of the PHP method_exists function in the main loops of the instruction processing and expression resolving methods, to find and call any extension methods added to implement features of a different language. At the moment it is still evolving, so as I developed the interpreters for the parser language and the template language, I’m making a decision whether any new features go into their own derived classes, or back into the base CoreLang class.