Record names of extracted files

A way is needed to record scanned file names for two purposes:

1. File names (and extensions) must be stored in the json metadata
properties recorded when using the --gen-json clamscan option. Future
work may use this to compare file extensions with detected file types.

2. File names are useful when interpretting tmp directory output when
using the --leave-temps option.

This commit enables file name retention for later use by storing file
names in the fmap header structure, if a file name exists.

To store the names in fmaps, an optional name argument has been added to
any internal scan API's that create fmaps and every call to these APIs
has been modified to pass a file name or NULL if a file name is not
required.  The zip and gpt parsers required some modification to record
file names.  The NSIS and XAR parsers fail to collect file names at all
and will require future work to support file name extraction.

Also:

- Added recursive extraction to the tmp directory when the
  --leave-temps option is enabled.  When not enabled, the tmp directory
  structure remains flat so as to prevent the likelihood of exceeding
  MAX_PATH.  The current tmp directory is stored in the scan context.

- Made the cli_scanfile() internal API non-static and added it to
  scanners.h so it would be accessible outside of scanners.c in order to
  remove code duplication within libmspack.c.

- Added function comments to scanners.h and matcher.h

- Converted a TDB-type macros and LSIG-type macros to enums for improved
  type safey.

- Converted more return status variables from `int` to `cl_error_t` for
  improved type safety, and corrected ooxml file typing functions so
  they use `cli_file_t` exclusively rather than mixing types with
  `cl_error_t`.

- Restructured the magic_scandesc() function to use goto's for error
  handling and removed the early_ret_from_magicscan() macro and
  magic_scandesc_cleanup() function.  This makes the code easier to
  read and made it easier to add the recursive tmp directory cleanup to
  magic_scandesc().

- Corrected zip, egg, rar filename extraction issues.

- Removed use of extra sub-directory layer for zip, egg, and rar file
  extraction.  For Zip, this also involved changing the extracted
  filenames to be randomly generated rather than using the "zip.###"
  file name scheme.
This commit is contained in:
Micah Snyder 2020-03-19 21:23:54 -04:00
parent 9f2de39e04
commit 005cbf5a37
67 changed files with 1054 additions and 722 deletions

View file

@ -1585,7 +1585,7 @@ int cli_scan_ole10(int fd, cli_ctx *ctx)
if (!read_uint32(fd, &object_size, FALSE))
return CL_CLEAN;
}
if (!(fullname = cli_gentemp(ctx ? ctx->engine->tmpdir : NULL))) {
if (!(fullname = cli_gentemp(ctx ? ctx->sub_tmpdir : NULL))) {
return CL_EMEM;
}
ofd = open(fullname, O_RDWR | O_CREAT | O_TRUNC | O_BINARY | O_EXCL,
@ -1598,7 +1598,7 @@ int cli_scan_ole10(int fd, cli_ctx *ctx)
cli_dbgmsg("cli_decode_ole_object: decoding to %s\n", fullname);
ole_copy_file_data(fd, ofd, object_size);
lseek(ofd, 0, SEEK_SET);
ret = cli_magic_scandesc(ofd, fullname, ctx);
ret = cli_magic_scandesc(ofd, fullname, ctx, NULL);
close(ofd);
if (ctx && !ctx->engine->keeptmp)
if (cli_unlink(fullname))
@ -1762,7 +1762,7 @@ cli_ppt_vba_read(int ifd, cli_ctx *ctx)
const char *ret;
/* Create a directory to store the extracted OLE2 objects */
dir = cli_gentemp(ctx ? ctx->engine->tmpdir : NULL);
dir = cli_gentemp(ctx ? ctx->sub_tmpdir : NULL);
if (dir == NULL)
return NULL;
if (mkdir(dir, 0700)) {