Software composition analysis on target source code
Abstract
A computer-implemented method (400) for performing software composition analysis of a target source code (90) for a computer program or a part thereof is disclosed herein. The method (400) involves performing (410) a first exploration process. The first exploration process comprises searching (412) a plurality of first software archives (10) originating from different sources in a global computer network (100) to find first occurrences (12) of the target source code (90) among source code files in the plurality of first software archives (10), and for every found first occurrence (12) of the target source code (90), collecting (414) a first set of key information (14) about matching source code files (16) or snippets (16a) therein. The method (400) further involves performing (420) a second exploration process. The second exploration process comprises searching (422) a plurality of second software archives (20) originating from one or more sources in the global computer network (100), the plurality of second software archives (20) being different from the plurality of first software archives (10), to find second occurrences (22) of the target source code (90) among source code snippets in the second software archives (20), and for every found second occurrence (22) of the target source code (90), collecting (424) a second set of key information (24) about matching source code snippets (26). The method (400) further comprises mapping (430) each matching source code snippet among the matching source code snippets (26) as collected in the second set of key information (24) to the matching source code files (16) or snippets (16a) therein as collected in the first set of key information (14), wherein the mapping (430) indicates whether an earlier version of said each matching source code snippet exists in the first set of key information (14). The method (400) further comprises, based on the mapped first set of key information (14) and second set of key information (24), determining (440) a software composition (92) of the target source code (90).
Claims
exact text as granted — not AI-modified1 . A computer-implemented method for performing software composition analysis of a target source code for a computer program or a part thereof, the method involving:
performing a first exploration process, comprising: searching a plurality of first software archives originating from different sources in a global computer network to find first occurrences of the target source code among source code files in the plurality of first software archives, and for every found first occurrence of the target source code, collecting a first set of key information about matching source code files or snippets therein; performing a second exploration process, comprising: searching a plurality of second software archives originating from one or more sources in the global computer network, the plurality of second software archives being different from the plurality of first software archives, to find second occurrences of the target source code among source code snippets in the second software archives, and for every found second occurrence of the target source code, collecting a second set of key information about matching source code snippets; mapping each matching source code snippet among the matching source code snippets as collected in the second set of key information to the matching source code files or snippets therein as collected in the first set of key information, wherein the mapping indicates whether an earlier version of said each matching source code snippet exists in the first set of key information; and based on the mapped first set of key information and second set of key information, determining a software composition of the target source code.
2 . The computer-implemented method according to claim 1 , wherein if an earlier version of said matching source code snippet exists in the first set of key information, the mapping further involves filtering said matching source code snippet from the second set of key information.
3 . The computer-implemented method according to claim 2 , wherein the filtering involves discarding, removing, hiding and/or down ranking said matching source code snippet from or in the second set of key information.
4 . The computer-implemented method according to claim 1 , wherein the software composition is indicative of origins, licenses, versions, vulnerabilities, comments, repositories, authors, file sizes, snippet sizes and/or resource locations associated with the target source code.
5 . The computer-implemented method according to claim 1 , wherein the plurality of first software archives are open source code archives.
6 . The computer-implemented method according to claim 1 , wherein the plurality of second software archives are Internet-based community-driven platform archives.
7 . The computer-implemented method according to claim 1 , wherein the plurality of first software archives and/or second software archives is any combination of open source code archives or Internet-based community-driven platform archives.
8 . The computer-implemented method according to claim 7 , wherein the open source code archives are one of or a combination of Github, Gitlab, or Bitbucket archives.
9 . The computer-implemented method according to claim 7 , wherein the Internet-based community-driven platform archives are one of or a combination of StackOverflow or other StackExchange platforms, Coderanch, Quora, Reddit, Google Groups, SitePoint, CodeProject, Google+Communities, Treehouse, Hacker News, DZone, Bytes, DaniWeb, Dream.In.Code, Pineapple, Lobsters, XDA Developers, CodeGuru, Programmers Heaven, FindNerd, Designers Talk, Hashnode or Mozzila Web Developer Community archives.
10 . The computer-implemented method according to claim 1 , wherein the method further involves creating a virtual file tree of the software composition of the target source code.
11 . The computer-implemented method according to claim 1 , wherein the method further involves an initial step of receiving a request for performing software composition analysis of the target source code
12 . The computer-implemented method according to claim 1 , wherein the method further involves compiling the software composition into a software composition analysis report, and returning said report.
13 . The computer-implemented method according to claim 1 , wherein the first set of key information comprises one or more keywords from a plurality of attributes of the matching source code file or snippets therein and/or the first software archive in which it was found.
14 . The computer-implemented method according to claim 1 , wherein the second set of key information comprises one or more keywords from a plurality of attributes of the matching source code snippets and/or the second software archive in which it was found.
15 . The computer-implemented method according to claim 14 , wherein the plurality of attributes includes at least two of the following: an author, a repository name, a filename and a resource location of the matching source code file and/or the snippets and/or the software archive in which it was found.
16 . The computer-implemented method according to claim 1 , wherein the plurality of first and/or second software archives originating from different sources in the global computer network are maintained in a local data repository.
17 . An apparatus for performing software composition analysis of a target source code for a computer program or a part thereof, the apparatus comprising a processing device being configured for:
performing a first exploration process, comprising: searching a plurality of first software archives from different sources in a global computer network to find first occurrences of the target source code among source code files in the plurality of first software archives, and for every found first occurrence of the target source code, collecting a first set of key information about matching source code files or snippets therein; performing a second exploration process, comprising: searching a plurality of second software archives from different sources in the global computer network, the plurality of second software archives being different from the plurality of first software archives, to find second occurrences of the target source code among snippets in the second software archives, and for every found second occurrence of the target source code, collecting a second set of key information about matching source code snippets; mapping each matching source code snippet among the matching source code snippets as collected in the second set of key information to the matching source code files or snippets therein as collected in the first set of key information, wherein the mapping indicates whether an earlier version of said each matching source code snippet exists in the first set of key information and based on the mapped first set of key information and second set of key information, determining a software composition of the target source code.
18 - 20 . (canceled)Join the waitlist — get patent alerts
Track US2022300277A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.