tensorflow / tensorflow/java

Converting TensorFlow markdown to JavaDoc text in op_generator

Offen
#213 5 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Vorherrschende Sprache
Java
Sterne
928
Forks
227
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

@karllessard @Craigacp

I have been experimenting with converting the TF Markdown text to JavaDoc format in the op_generator code.
I did this by creating another c++ class, that calls out to Python using the Python C library. This runs the Python marko package with my own marko renderer class javadoc_renderer.JavaDocRenderer that converts markdown to JavaDoc.
In the C++ class, SourceWriter, I call out to the python code to convert the Markdown text to JavaDoc. The converted JavaDoc code is then written out to the class.

Here is an example of the old and new generated JavaDoc for org.tensorflow.op.math.Abs:

Current JavaDoc:

/**
 * Computes the absolute value of a tensor.
 * <p>
 * Given a tensor `x`, this operation returns a tensor containing the absolute
 * value of each element in `x`. For example, if x is an input element and y is
 * an output element, this operation computes \\(y = |x|\\).
 * 
 * @param <T> data type for {@code y()} output
 */

New JavaDoc:

/**
 * <p>Computes the absolute value of a tensor.</p>
 * <p>
 * <p>Given a tensor <code>x</code>, this operation returns a tensor containing the absolute
 * value of each element in <code>x</code>. For example, if x is an input element and y is
 * an output element, this operation computes \(y = |x|\).</p>
 * 
 * @param <T> data type for {@code y()} output
 */

There still needs some tweaks to JavaDoc output, like <p> on a single line.
Also, I am still chasing down an infrequent error where the conversion string gets garbled.

I have made several design decision that should probably be discussed. For example, I put my Python module in bazel-bin and point the PYTHONPATH to it in build.sh.

env PYTHONPATH=:$BAZEL_BIN/markdown_javadoc $BAZEL_BIN/java_op_generator \
    --output_dir=$GEN_SRCS_DIR \
    --api_dirs=$BAZEL_SRCS/external/org_tensorflow/tensorflow/core/api_def/base_api,src/bazel/api_def \
    $TENSORFLOW_LIB

Also, I cannot figure out how to bring in the python library from the framework into the BUILD file.
For now, I have it hard coded.

tf_cc_binary(
    name = "java_op_generator",
    linkopts = select({
        "@org_tensorflow//tensorflow:windows": [],
        "//conditions:default": [
            "-lm",
            "-L/Library/Frameworks/Python.framework/Versions/3.7/lib/python3.7/config-3.7m-darwin",
            "-lpython3.7"
          ],
        }),
    deps = [
        ":java_op_gen_lib",
    ],
)

Any help on setting the bazel rules for include the python library would be appreciated.

I did find @org_tensorflow//third_party/python_runtime:headers, which I added as a dependency in the cc_library section of BUILD. This allowed me to compile the c++ code with the Python.h header.

cc_library(
    name = "java_op_gen_lib",
    srcs = [
        "src/bazel/op_generator/op_gen_main.cc",
        "src/bazel/op_generator/op_generator.cc",
        "src/bazel/op_generator/op_specs.cc",
        "src/bazel/op_generator/source_writer.cc",
        "src/bazel/op_generator/markdown_javadoc.cc",
    ],
    hdrs = [
        "src/bazel/op_generator/java_defs.h",
        "src/bazel/op_generator/op_generator.h",
        "src/bazel/op_generator/op_specs.h",
        "src/bazel/op_generator/source_writer.h",
        "src/bazel/op_generator/markdown_javadoc.h",
    ],
    copts = tf_copts(),
    deps = [
        "@org_tensorflow//tensorflow/core:framework",
        "@org_tensorflow//tensorflow/core:lib",
        "@org_tensorflow//tensorflow/core:op_gen_lib",
        "@org_tensorflow//tensorflow/core:protos_all_cc",
        "@org_tensorflow//third_party/python_runtime:headers",
        "@com_googlesource_code_re2//:re2",
    ],
)

I can create a draft PR if you want to look at the whole project, so we can iterate on some of the design decisions, and figure out how to link with the Python C library in a bazel friendly way.

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne mit dem Lesen von src/bazel/op_generator/source_writer.cc und markdown_javadoc.cc und untersuche anschließend die BUILD- und build.sh-Ausschnitte für die Python-Integration und den generierten JavaDoc-Pfad. Für den Abschluss wären eine abgestimmte Bazel-freundliche Python-Verknüpfung, eine korrigierte JavaDoc-Ausgabe und die Behebung der hier gemeldeten gelegentlich fehlerhaften Konvertierung erforderlich.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
cpp, java, python, tensorflow
Bereich
build-system, documentation, tooling
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.