CEA presented 1 platform and 9 posters at SETAC Seville in May 2024. We will be showcasing each of these presentations in a series of ‘SETAC Spotlight’ articles. This week it is:
SETAC Spotlight: Can custom-written codes be GLP compliant?
Authors: Marie Brown(1), Hanna Schuster(2) and Bashir Surfraz(3)
Cambridge Environmental Assessments, RSK ADAS, Boxworth, UK
Statistical programming tools such as SAS, Matlab, Python and R have revolutionized data analysis processes, particularly for large datasets. Additionally, KNIME (Konstanz Information Miner), a popular low-coding, free, analytic platform with R and Python integration, offers great future potential. The free open-source system R has become a powerful tool by providing greater flexibility and innovation in communicating data by using various packages with different functionalities. R packages are constantly updated, which is great for statistical advancements and improvements, but is problematic for use in and validation for Good Laboratory Practice (GLP) studies. Validation for GLP compliance requires proof that the program is operating correctly, a risk assessment to determine the probability that errors or misuse could occur, and documentation and evidence of change and version control. The outcome is a high level of assurance that the program and functions are fit for purpose and results are reliable and trustworthy. Whilst there are documents setting out coding standards and best practices to ensure that the data processing and analysis can be clearly followed by another user, this does not fulfil the requirements of GLP validation. There is a lack of clarity and guidance on how such customised tools e.g., codes/workflows, can be validated for GLP compliance. Although OECD guidance on computerised systems has been revised to address more sophisticated data analysis, customised systems are classified as the highest risk level. Thus, the validation process poses a significant challenge. Nonetheless, there are major advantages of coding and building data processing pipelines regarding the transparency of data handling and analysis, particularly when using open-source software such as R and KNIME. An R script or a KNIME workflow can be annotated to explain each data handling step and it can be re-run by any user on any computer, resulting in high reproducibility of the data journey. Using KNIME, together with coding software, increases the readability and understanding of each step (node) within a workflow and brings together separate data silos (loading, processing, and analyses). In this modular approach, each node or workflow can be individually validated and archived.
This poster provides a potential solution on how to validate programming tools within KNIME and aims to open up discussions in the SETAC and wider scientific and GLP communities.

If you would like to discuss any of the topics raised in this article, feel free to contact us.
This poster is available for free download.
You can find all of the other posters that CEA presented at SETAC Seville here. You can also find all of our publications from previous conferences and links to journal articles we have authored on our library page.
Enjoyed this article? Avoid missing out on our future news articles by signing up – you can select which topics you are interested in and can unsubscribe at any time.
