DIS3GNO

A visual environment for designing and running data mining workflows in the knowledge grid Data mining tasks are often composed by multiple stages that may be linked to each other to form various execution flows. Moreover, data mining tasks are often distributed since they involve data and tools located over geographically distributed environments, like the Grid. Therefore, it is fundamental to exploit effective formalisms, such as workflows, to model data mining tasks that are both multi-staged and distributed. The goal of this work is defining a workflow formalism and providing a visual software environment, named DIS3GNO, to design and execute distributed data mining tasks over the knowledge grid, a service-oriented framework for distributed data mining on the grid. DIS3GNO supports all the phases of a distributed data mining task, including composition, execution, and results visualization. The paper provides a description of DIS3GNO, some relevant use cases implemented by it, and a performance evaluation of the system.