Recommending Relevant Classes for Infrequent API Classes
Project summary As APIs are many and complicated, it is difficult to learn their usages, and programmers can introduce API-related bugs in their code. To assist programming and reduce API-related bugs, in the literature, researchers have proposed various approaches that assist programmin
Project summary
As APIs are many and complicated, it is difficult to learn their usages, and programmers can introduce API-related bugs in their code. To assist programming and reduce API-related bugs, in the literature, researchers have proposed various approaches that assist programming with APIs, and many approaches need API patterns that are mined from clients or documents. However, it is rather challenging to mine complete API patterns. The incompleteness of mined patterns is a barrier for many API-related research. For example, if a tool determines that the violations of known API patterns are bugs, the tool can produce many false alarms, when API patterns are incomplete.
Our study shows that from clients, it is infeasible to mine frequent patterns for many APIs, especially when their versions are considered. From documents, many API patterns are not explicitly described either. As a result, it is quite challenging to improve the completeness of API patterns.
In this project, we propose a novel research direction to improve the completeness of API patterns, and it can mine patterns that do not appear in clients or documents. From clients, it is typically able to mine only limited API patterns. Our idea is to generate training data from mined patterns and documents, and to learn a model that can predict more infrequent patterns. To show the feasibility of this research direction, we take a classical type of API patterns, relevant APIs, as an example. For an API class, an relevant API is a set of API classes that are called with this API class. The prior approaches mine relevant APIs with association mining. From mined frequent call sets, our tool, APIRel, generates positive instances and negative instances, and extract their features by analyzing API documents. In this way, it trains models that can predict infrequent relevant APIs that do not appear in known clients.
Setting
In our evaluation, we select the APIs of accumulo, cassandra, karaf, lucene, and poi. Their API documents are listed as follows:
accumulo: https://tohidemyname.github.io/accumulodoc/
cassandra: https://tohidemyname.github.io/cassandradoc/
karaf: https://tohidemyname.github.io/karafdoc/
lucene: https://tohidemyname.github.io/lucenedoc/
poi: https://tohidemyname.github.io/poidoc/
Artifacts Overview
manual_sample contains about 700 manually labeled API class pairs (about 140 pairs per library across five Apache projects). Each pair was inspected by reading the official API documentation of both classes and checking whether they can be called together to implement a specific functionality.
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.